CatEncoder#

class skrub.CatEncoder(max_categories=10)[source]#

Encode a single categorical column combining OneHotEncoder and TargetEncoder.

A note on using single column transformations

CatEncoder is a type of single column transformation . Unlike most scikit-learn estimators, its fit, transform and fit_transform methods expect a single column (e.g. Series) not a full dataframe. To apply this transformer to one or more columns in a dataframe, use it in a ApplyToCols or a TableVectorizer.

To apply to all columns:

ApplyToCols(CatEncoder())

To apply to selected columns:

ApplyToCols(CatEncoder(), cols=['col_name_1', 'col_name_2'])

This transformer applies a OneHotEncoder to encode frequent categories into binary one-hot columns, and a TargetEncoder to target-encode the column.

Parameters:
max_categoriespython:int or python:None, default=10

Maximum number of categories for the OneHotEncoder. If there are more categories, the remaining ones are grouped into an infrequent category.

Attributes:
one_hot_encoder_OneHotEncoder

The fitted OneHotEncoder instance.

target_encoder_TargetEncoder

The fitted TargetEncoder instance.

one_hot_outputs_python:list of python:str

Feature names created by the one-hot encoder.

target_outputs_python:list of python:str

Feature names using the _target_sklearn suffix (followed by the class label for multiclass targets). Collisions with one-hot features are resolved by adding a random __skrub_<token>__ suffix.

all_outputs_python:list of python:str

The list of feature names created by the transformer.

Examples

>>> import pandas as pd
>>> from skrub import CatEncoder
>>> s = pd.Series(["a", "b", "c", "d", "e"] * 4, name="col")
>>> y = pd.Series([1, 0, 1, 0, 1] * 4)
>>> enc = CatEncoder(max_categories=3)
>>> enc.fit_transform(s, y).head(2)
   col_d  col_e  col_infrequent_sklearn  col_target_sklearn
0    0.0    0.0                     1.0                 ...
1    0.0    0.0                     1.0                 ...

Methods

fit(column, y)

Fit the encoder to a categorical column.

fit_transform(column, y)

Fit the encoder and transform a categorical column.

get_feature_names_out([input_features])

Return the names of all generated output features.

get_params([deep])

Get parameters for this estimator.

set_output(*[, transform])

Set output container.

set_params(**params)

Set the parameters of this estimator.

set_transform_request(*[, column])

Configure whether metadata should be requested to be passed to the transform method.

transform(column)

Transform a single column using fitted OneHotEncoder and TargetEncoder.

fit(column, y)[source]#

Fit the encoder to a categorical column.

Parameters:
columnPandas or Polars Series

The single column to fit.

yPandas or Polars Series, DataFrame, or array_like

Target values for target encoding.

Returns:
self

The fitted encoder.

fit_transform(column, y)[source]#

Fit the encoder and transform a categorical column.

Parameters:
columnPandas or Polars Series

The single column to transform.

yPandas or Polars Series, DataFrame, or array_like

Target values for target encoding.

Returns:
res_dfPandas or Polars DataFrame

DataFrame containing one-hot and target-encoded features.

get_feature_names_out(input_features=None)[source]#

Return the names of all generated output features.

Parameters:
input_featuresarray_like of python:str or python:None, default=None

Ignored.

Returns:
python:list of python:str

Feature names generated by the encoder.

get_params(deep=True)[source]#

Get parameters for this estimator.

Parameters:
deeppython:bool, default=True

If True, will return the parameters for this estimator and contained subobjects that are estimators.

Returns:
paramspython:dict

Parameter names mapped to their values.

set_output(*, transform=None)[source]#

Set output container.

Refer to the user guide for more details and Introducing the set_output API for an example on how to use the API.

Parameters:
transform{“default”, “pandas”, “polars”}, default=None

Configure output of transform and fit_transform.

  • "default": Default output format of a transformer

  • "pandas": DataFrame output

  • "polars": Polars output

  • None: Transform configuration is unchanged

Added in version 1.4: "polars" option was added.

Returns:
selfestimator instance

Estimator instance.

set_params(**params)[source]#

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects (such as Pipeline). The latter have parameters of the form <component>__<parameter> so that it’s possible to update each component of a nested object.

Parameters:
**paramspython:dict

Estimator parameters.

Returns:
selfestimator instance

Estimator instance.

set_transform_request(*, column='$UNCHANGED$')[source]#

Configure whether metadata should be requested to be passed to the transform method.

Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()). Please check the User Guide on how the routing mechanism works.

The options for each parameter are:

  • True: metadata is requested, and passed to transform if provided. The request is ignored if metadata is not provided.

  • False: metadata is not requested and the meta-estimator will not pass it to transform.

  • None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.

  • str: metadata should be passed to the meta-estimator with this given alias instead of the original name.

The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.

Added in version 1.3.

Parameters:
columnpython:str, python:True, python:False, or python:None, default=sklearn.utils.metadata_routing.UNCHANGED

Metadata routing for column parameter in transform.

Returns:
selfobject

The updated object.

transform(column)[source]#

Transform a single column using fitted OneHotEncoder and TargetEncoder.

Parameters:
columnPandas or Polars Series

The column to transform.

Returns:
res_dfPandas or Polars DataFrame

Transformed features.