ToCategorical#

class skrub.ToCategorical(accept_int=False)[source]#

Convert a string column to Categorical dtype.

A note on using single column transformations

ToCategorical is a type of single column transformation . Unlike most scikit-learn estimators, its fit, transform and fit_transform methods expect a single column (e.g. Series) not a full dataframe. To apply this transformer to one or more columns in a dataframe, use it in a ApplyToCols or a TableVectorizer.

To apply to all columns:

ApplyToCols(ToCategorical())

To apply to selected columns:

ApplyToCols(ToCategorical(), cols=['col_name_1', 'col_name_2'])

This transformer ensures that a given string or categorical column has Categorical dtype so that it is treated as categorical by downstream transformers and learners.

Parameters:
accept_intpython:bool, default=False

How to handle numeric columns. If False, no numeric columns will be accepted. If True, will convert integer columns to categorical.

Notes

The main benefit of converting columns to categorical is that categorical columns can be recognized by scikit-learn’s HistGradientBoostingRegressor and HistGradientBoostingClassifier with their categorical_features='from_dtype' option. This transformer is therefore particularly useful as the low_cardinality_transformer parameter of the TableVectorizer when combined with one of those supervised learners.

A pandas column with dtype string or object containing strings, or a polars column with dtype String, is converted to a categorical column. Categorical columns are passed through.

If accept_int is set to True, then integer columns are also accepted and converted to categorical. The default value is False.

Any other type of column is rejected by raising a RejectColumn exception. Note: the TableVectorizer only sends string or categorical columns to its low_cardinality_transformer, regardless of the inputted value of accept_int. Therefore it is always safe to use a ToCategorical instance as the low_cardinality_transformer.

The output of transform also always has a Categorical dtype. The categories are not necessarily the same across different calls to transform. Indeed, scikit-learn estimators do not inspect the dtype’s categories but the actual values. Converting to a Categorical is therefore just a way to mark a column and indicate to downstream estimators that this column should be treated as categorical. Ensuring they are encoded consistently, handling unseen categories at test time, etc. is the responsibility of encoders such as OneHotEncoder and LabelEncoder, or of estimators that handle categories themselves such as HistGradientBoostingRegressor.

Examples

>>> import pandas as pd
>>> from skrub import ToCategorical

A string column is converted to a categorical column.

>>> s = pd.Series(['one', 'two', None], name='c')
>>> s
0     one
1     two
2     ...
Name: c, dtype: ...
>>> to_cat = ToCategorical()
>>> to_cat.fit_transform(s)
0    one
1    two
2    ...
Name: c, dtype: ...
Categories (2, ...): ['one', 'two']

The dtypes (the list of categories) of the outputs of transform may vary. This transformer only ensures the dtype is Categorical to mark the column as such for downstream encoders which will perform the actual encoding.

>>> s = pd.Series(['four', 'five'], name='c')
>>> to_cat.transform(s)
0    four
1    five
Name: c, dtype: category
Categories (2, ...): ['five', 'four']

Columns that already have a Categorical dtype are passed through:

>>> s = pd.Series(['one', 'two'], name='c', dtype='category')
>>> to_cat.fit_transform(s) is s
True

Columns that are not strings nor categorical are rejected:

>>> to_cat.fit_transform(pd.Series([1.1, 2.2], name='c'))
Traceback (most recent call last):
    ...
skrub.core.RejectColumn: Column 'c' does not contain only strings...

Unless accept_int is set to True, in which case integer columns are accepted:

>>> to_cat = ToCategorical(accept_int=True)
>>> to_cat.fit_transform(pd.Series([1, 2], name='c'))
0    1
1    2
Name: c, dtype: category
Categories (2, int64): [1, 2]

object columns that do not contain only strings are also rejected:

>>> s = pd.Series(['one', 1], name='c')
>>> to_cat.fit_transform(s)
Traceback (most recent call last):
    ...
skrub.core.RejectColumn: Column 'c' does not contain only strings...

No special handling of StringDtype vs object columns is done, the behavior is the same as pd.astype('category'): if the input uses the extension dtype, the categories of the output will, too.

>>> s = pd.Series(['cat A', 'cat B', None], name='c', dtype='string')
>>> s
0    cat A
1    cat B
2     <NA>
Name: c, dtype: string
>>> to_cat.fit_transform(s)
0    cat A
1    cat B
2     <NA>
Name: c, dtype: category
Categories (2, string): [...]
>>> _.cat.categories.dtype

Polars string columns are converted to the Categorical dtype (not Enum). As for pandas, categories may vary across calls to transform.

>>> import pytest
>>> pl = pytest.importorskip("polars")
>>> s = pl.Series('c', ['one', 'two', None])
>>> to_cat.fit_transform(s)
shape: (3,)
Series: 'c' [cat]
[
    "one"
    "two"
    null
]

Polars Categorical or Enum columns are passed through:

>>> s = pl.Series('c', ['one', 'two'], dtype=pl.Enum(['one', 'two', 'three']))
>>> s
shape: (2,)
Series: 'c' [enum]
[
    "one"
    "two"
]
>>> to_cat.fit_transform(s) is s
True

Methods

fit(column[, y])

Fit the transformer.

fit_transform(column[, y])

Fit the encoder and transform a column.

get_feature_names_out([input_features])

Get the output feature names.

get_params([deep])

Get parameters for this estimator.

set_output(*[, transform])

Default no-op implementation for set_output.

set_params(**params)

Set the parameters of this estimator.

set_transform_request(*[, column])

Configure whether metadata should be requested to be passed to the transform method.

transform(column)

Transform a column.

fit(column, y=None, **kwargs)[source]#

Fit the transformer.

This default implementation simply calls fit_transform() and returns self.

Parameters:
columna pandas or polars Series

Unlike most scikit-learn transformers, single-column transformers transform a single column, not a whole dataframe.

ycolumn or dataframe

Prediction targets.

**kwargs

Extra named arguments are passed to self.fit_transform().

Returns:
self

The fitted transformer.

fit_transform(column, y=None)[source]#

Fit the encoder and transform a column.

Parameters:
columnpandas or polars Series
ypython:None

Ignored.

Returns:
transformedpandas or polars Series

The input transformed to Categorical.

get_feature_names_out(input_features=None)[source]#

Get the output feature names.

Parameters:
input_featuresarray_like of python:str, default=None

Input feature names. Ignored.

Returns:
python:list of python:str

The names of the output features.

get_params(deep=True)[source]#

Get parameters for this estimator.

Parameters:
deeppython:bool, default=True

If True, will return the parameters for this estimator and contained subobjects that are estimators.

Returns:
paramspython:dict

Parameter names mapped to their values.

set_output(*, transform=None)[source]#

Default no-op implementation for set_output.

Skrub transformers already output dataframes of the correct type by default so there is usually no need for set_output to do anything.

Subclasses are of course free to redefine set_output (e.g. by inheriting from TransformerMixin before SingleColumnTransformer).

Parameters:
transformpython:str or python:None, default=None

Ignored.

Returns:
SingleColumnTransformer

Returns self.

set_params(**params)[source]#

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects (such as Pipeline). The latter have parameters of the form <component>__<parameter> so that it’s possible to update each component of a nested object.

Parameters:
**paramspython:dict

Estimator parameters.

Returns:
selfestimator instance

Estimator instance.

set_transform_request(*, column='$UNCHANGED$')[source]#

Configure whether metadata should be requested to be passed to the transform method.

Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()). Please check the User Guide on how the routing mechanism works.

The options for each parameter are:

  • True: metadata is requested, and passed to transform if provided. The request is ignored if metadata is not provided.

  • False: metadata is not requested and the meta-estimator will not pass it to transform.

  • None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.

  • str: metadata should be passed to the meta-estimator with this given alias instead of the original name.

The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.

Added in version 1.3.

Parameters:
columnpython:str, python:True, python:False, or python:None, default=sklearn.utils.metadata_routing.UNCHANGED

Metadata routing for column parameter in transform.

Returns:
selfobject

The updated object.

transform(column)[source]#

Transform a column.

Parameters:
columnpandas or polars Series

The input to transform.

Returns:
transformedpandas or polars Series

The input transformed to Categorical.