Caching for faster recomputation#
The results of estimators added to a DataOp with .skb.apply() and of functions added with deferred() or
.skb.apply_func() can be cached.
This is a fairly experimental feature. The feature and its API may change to improve the identification of identical calls, the management of the cache directory, etc. If you experiment with it, please provide feedback!
Caching can save a lot of computation when we run the same operation again. This typically happens when some step in a pipeline has changed (during hyperparameter search or because we modified our pipeline), but some earlier steps remain the same and their results can be reused. For example suppose we have evaluated the following DataOp:
skrub.X().skb.apply(skrub.TableVectorizer()).skb.apply(
RandomForestRegressor(), y=skrub.y()
)
and later we modify it to replace the final estimator and run it on the same data:
skrub.X().skb.apply(skrub.TableVectorizer()).skb.apply(
HistGradientBoostingRegressor(), y=skrub.y()
)
The transformations applied by the TableVectorizer() can be reused.
For this to happen, we need to enable caching in the configuration, either
skrub.set_config(cache=True) to store cached results in a default location
(a _cache subdirectory inside the data dir get_config()["data_dir"])
or skrub.set_config(cache="/path/to/cache_dir/") to specify where to store
the cache.
Note
The caching mechanism discussed here is about persisting results on disk across different evaluations of a DataOp, or evaluations of different DataOps. Retaining intermediate results that are used in several places in a single DataOp in-memory until they are no longer needed, during a single evaluation of the DataOp, always happens.
Forbidding caching for specific nodes#
When adding nodes to a DataOp, we can specify that their results should never be
cached, even when caching is enabled in the configuration. This is useful if
caching causes errors (e.g. because the arguments or result cannot be
serialized), if result is not deterministic and should be recomputed every time
(for example fetching some information from the network), or if we know that the
function is very fast and caching hinders performance instead of improving it.
This is achieved by passing no_cache=True to deferred(),
.skb.apply() or .skb.apply_func().
Limiting the cache directory size#
By default skrub will start a subprocess to prune the cache in order to limit
its size, once per python program using the skrub cache. This can be controlled
with the target_cache_size config option. Set it to
None to disable this pruning altogether, to an int for a size in bytes, or
to a string like '3K', '3M', '3G'. The default is '2G'.
Note this is a rough target size and not a strict limit. In particular, as the pruning only runs once (the first time the caching is used in a program), the cache may grow afterwards and become bigger than the target size.
The cache can be stale#
Skrub relies on joblib for caching. Some effort is done to detect if the
code of cached functions has changed (thus invalidating cached results), but
this is on a best-effort basis and is fairly brittle. For example, changes to a
helper that is called by the cached function are not detected. Changes to the
code of estimators passed to DataOp.skb.apply() are not detected. (Note
the same limitations apply to joblib in general, and for example the memory
parameter of scikit-learn pipelines).
To handle this consider complementing the automated heuristics in place with
some manual intervention, such as applying no_cache=True for functions or
estimators that you are actively modifying, manually deleting the cache
directory when appropriate, or enabling caching only in experimental
environments.