set_config#

skrub.set_config(use_table_report_data_ops=None, table_report_plots_threshold=None, table_report_associations_threshold=None, table_report_n_rows=None, table_report_verbosity=None, subsampling_seed=None, max_plot_columns=None, max_association_columns=None, enable_subsampling=None, float_precision=None, cardinality_threshold=None, data_dir=None, cache=unchanged, target_cache_size=unchanged, eager_data_ops=None, data_ops_open_graph_dropdown=None)[source]#

Set global skrub configuration.

Parameters:
use_table_report_data_opspython:bool, default=None

The type of HTML representation used for the dataframes preview in skrub DataOps. If None, falls back to the current configuration, which is True by default.

  • If True, TableReport will be used.

  • If False, the original Pandas or Polars dataframe display will be used.

This configuration can also be set with the SKB_USE_TABLE_REPORT_DATA_OPS environment variable.

table_report_plots_thresholdpython:int, default=None

Maximum number of columns for which distribution plots are generated in TableReport when plot_distributions="auto" (the default). Dataframes with more columns will skip plots. Default is 30.

This configuration can also be set with the SKB_TABLE_REPORT_PLOTS_THRESHOLD environment variable.

table_report_associations_thresholdpython:int, default=None

Maximum number of columns for which associations are computed in TableReport when compute_associations="auto" (the default). Dataframes with more columns will skip associations. Default is 30.

This configuration can also be set with the SKB_TABLE_REPORT_ASSOCIATIONS_THRESHOLD environment variable.

table_report_n_rowspython:int, default=None

Set the default number of rows displayed in TableReport when the n_rows parameter is not explicitly passed. Default is 10.

This configuration can also be set with the SKB_TABLE_REPORT_N_ROWS environment variable.

table_report_verbositypython:int, default=None

Set the level of verbosity of the TableReport. Default is 1 (print the progress bar). Refer to the TableReport documentation for more details.

subsampling_seedpython:int, default=None

Set the random seed of subsampling in skrub DataOps skrub.DataOp.skb.subsample(), when how="random" is passed.

This configuration can also be set with the SKB_SUBSAMPLING_SEED environment variable.

enable_subsampling{‘default’, ‘disable’, ‘force’}, default=None

Control the activation of subsampling in skrub DataOps skrub.DataOp.skb.subsample(). Default is "default".

  • If "default", the behavior of skrub.DataOp.skb.subsample() is used.

  • If "disable", subsampling is never used, so skb.subsample becomes a no-op.

  • If "force", subsampling is used in all DataOps evaluation modes (eval(), fit_transform, etc.).

This configuration can also be set with the SKB_ENABLE_SUBSAMPLING environment variable.

float_precisionpython:int, default=3

Control the number of significant digits shown when formatting floats. Applies overall precision rather than fixed decimal places. Default is 3.

This configuration can also be set with the SKB_FLOAT_PRECISION environment variable.

cardinality_thresholdpython:int, default=40

Set the cardinality_threshold argument of TableVectorizer. Control the threshold value used to warn user if they have high cardinality columns in there dataset.

This configuration can also be set with the SKB_CARDINALITY_THRESHOLD environment variable.

data_dirpython:str or pathlib.Path, default=None

Set the data directory path for skrub datasets. If None, falls back to the current configuration.

  • If the SKB_DATA_DIRECTORY environment variable is set to an absolute path, that path will be used.

  • Otherwise, the default is ~/skrub_data.

This configuration can also be set with the SKB_DATA_DIRECTORY environment variable. The deprecated SKRUB_DATA_DIRECTORY is still supported with a deprecation warning.

cachepython:bool or python:str, default=False

Caching to use for the evaluation of DataOps.

  • If False (or None), no caching is used.

  • If True, the cache directory is in a default location (data_dir / _cache).

  • If a string (or Path), this path is used as the cache directory.

This configuration can also be set with the SKB_CACHE environment variable.

See Caching for faster recomputation for more information about caching.

target_cache_sizepython:int, python:str or python:None

When using caching, the cache will be pruned if it grows bigger than this size. Can be an int to set a size in bytes, or a string like ‘3000000K’, ‘3000M’, ‘3G’. None means never prune the cache. By default it is set to 2G; set to None if you prefer to manage the cache yourself.

This configuration can also be set with the SKB_TARGET_CACHE_SIZE environment variable (Set it to ‘None’ or ‘’ to mean None ie no limit).

See Caching for faster recomputation for more information about caching.

eager_data_opspython:bool, default=True

Eagerly perform checks on the DataOps as soon they are created, and compute previews if preview data is available. If disabled, those checks are delayed until the DataOp is actually used (e.g. by calling .skb.eval() or make_learner()), and previews are not computed.

This option is used to speed-up the creation of large DataOps containing many nodes. It can also be useful in rare cases where a DataOp needs no inputs (for example it relies on a hard-coded filename to load data) but we want to prevent it from computing preview results as soon as it is constructed and delay computation until we explicitly request it. For most DataOps that do need inputs (contain skrub.var() nodes), previews can also be disabled simply by not providing preview data to skrub.var().

This configuration can also be set with the SKB_EAGER_DATA_OPS environment variable.

data_ops_open_graph_dropdownpython:bool, default=False

When displaying a DataOp that has a preview value in a jupyter notebook, should the dropdown that reveals the computational graph drawing be open (if True) or close (if False). This option mostly exists to control the display of DataOps in the skrub documentation examples. This configuration can also be set with the SKB_DATA_OPS_OPEN_GRAPH_DROPDOWN environment variable.

See also

get_config

Retrieve current values for global configuration.

config_context

Context manager for global skrub configuration.

Examples

>>> from skrub import set_config
>>> set_config(use_table_report_data_ops=True)