CatEncoder#
- class skrub.CatEncoder(max_categories=10)[source]#
Encode a single categorical column combining OneHotEncoder and TargetEncoder.
A note on using single column transformations
CatEncoderis a type of single column transformation . Unlike most scikit-learn estimators, itsfit,transformandfit_transformmethods expect a single column (e.g. Series) not a full dataframe. To apply this transformer to one or more columns in a dataframe, use it in aApplyToColsor aTableVectorizer.To apply to all columns:
ApplyToCols(CatEncoder())
To apply to selected columns:
ApplyToCols(CatEncoder(), cols=['col_name_1', 'col_name_2'])
This transformer applies a
OneHotEncoderto encode frequent categories into binary one-hot columns, and aTargetEncoderto target-encode the column.- Parameters:
- max_categories
python:intorpython:None, default=10 Maximum number of categories for the
OneHotEncoder. If there are more categories, the remaining ones are grouped into an infrequent category.
- max_categories
- Attributes:
- one_hot_encoder_
OneHotEncoder The fitted
OneHotEncoderinstance.- target_encoder_TargetEncoder
The fitted
TargetEncoderinstance.- one_hot_outputs_
python:listofpython:str Feature names created by the one-hot encoder.
- target_outputs_
python:listofpython:str Feature names using the
_target_sklearnsuffix (followed by the class label for multiclass targets). Collisions with one-hot features are resolved by adding a random__skrub_<token>__suffix.- all_outputs_
python:listofpython:str The list of feature names created by the transformer.
- one_hot_encoder_
Examples
>>> import pandas as pd >>> from skrub import CatEncoder >>> s = pd.Series(["a", "b", "c", "d", "e"] * 4, name="col") >>> y = pd.Series([1, 0, 1, 0, 1] * 4) >>> enc = CatEncoder(max_categories=3) >>> enc.fit_transform(s, y).head(2) col_d col_e col_infrequent_sklearn col_target_sklearn 0 0.0 0.0 1.0 ... 1 0.0 0.0 1.0 ...
Methods
fit(column, y)Fit the encoder to a categorical column.
fit_transform(column, y)Fit the encoder and transform a categorical column.
get_feature_names_out([input_features])Return the names of all generated output features.
get_params([deep])Get parameters for this estimator.
set_output(*[, transform])Set output container.
set_params(**params)Set the parameters of this estimator.
set_transform_request(*[, column])Configure whether metadata should be requested to be passed to the
transformmethod.transform(column)Transform a single column using fitted OneHotEncoder and TargetEncoder.
- fit(column, y)[source]#
Fit the encoder to a categorical column.
- Parameters:
- columnPandas or Polars
Series The single column to fit.
- yPandas or Polars
Series, DataFrame, or array_like Target values for target encoding.
- columnPandas or Polars
- Returns:
- self
The fitted encoder.
- fit_transform(column, y)[source]#
Fit the encoder and transform a categorical column.
- Parameters:
- columnPandas or Polars
Series The single column to transform.
- yPandas or Polars
Series, DataFrame, or array_like Target values for target encoding.
- columnPandas or Polars
- Returns:
- res_dfPandas or Polars DataFrame
DataFrame containing one-hot and target-encoded features.
- get_feature_names_out(input_features=None)[source]#
Return the names of all generated output features.
- Parameters:
- input_featuresarray_like of
python:strorpython:None, default=None Ignored.
- input_featuresarray_like of
- Returns:
python:listofpython:strFeature names generated by the encoder.
- get_params(deep=True)[source]#
Get parameters for this estimator.
- Parameters:
- deep
python:bool, default=True If True, will return the parameters for this estimator and contained subobjects that are estimators.
- deep
- Returns:
- params
python:dict Parameter names mapped to their values.
- params
- set_output(*, transform=None)[source]#
Set output container.
Refer to the user guide for more details and Introducing the set_output API for an example on how to use the API.
- Parameters:
- transform{“default”, “pandas”, “polars”}, default=None
Configure output of
transformandfit_transform."default": Default output format of a transformer"pandas": DataFrame output"polars": Polars outputNone: Transform configuration is unchanged
Added in version 1.4:
"polars"option was added.
- Returns:
- selfestimator instance
Estimator instance.
- set_params(**params)[source]#
Set the parameters of this estimator.
The method works on simple estimators as well as on nested objects (such as
Pipeline). The latter have parameters of the form<component>__<parameter>so that it’s possible to update each component of a nested object.- Parameters:
- **params
python:dict Estimator parameters.
- **params
- Returns:
- selfestimator instance
Estimator instance.
- set_transform_request(*, column='$UNCHANGED$')[source]#
Configure whether metadata should be requested to be passed to the
transformmethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed totransformif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it totransform.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.
- Parameters:
- column
python:str,python:True,python:False, orpython:None, default=sklearn.utils.metadata_routing.UNCHANGED Metadata routing for
columnparameter intransform.
- column
- Returns:
- selfobject
The updated object.