ToCategorical#
- class skrub.ToCategorical(accept_int=False)[source]#
Convert a string column to Categorical dtype.
A note on using single column transformations
ToCategoricalis a type of single column transformation . Unlike most scikit-learn estimators, itsfit,transformandfit_transformmethods expect a single column (e.g. Series) not a full dataframe. To apply this transformer to one or more columns in a dataframe, use it in aApplyToColsor aTableVectorizer.To apply to all columns:
ApplyToCols(ToCategorical())
To apply to selected columns:
ApplyToCols(ToCategorical(), cols=['col_name_1', 'col_name_2'])
This transformer ensures that a given string or categorical column has Categorical dtype so that it is treated as categorical by downstream transformers and learners.
- Parameters:
- accept_int
python:bool, default=False How to handle numeric columns. If
False, no numeric columns will be accepted. IfTrue, will convert integer columns to categorical.
- accept_int
Notes
The main benefit of converting columns to categorical is that categorical columns can be recognized by scikit-learn’s
HistGradientBoostingRegressorandHistGradientBoostingClassifierwith theircategorical_features='from_dtype'option. This transformer is therefore particularly useful as thelow_cardinality_transformerparameter of theTableVectorizerwhen combined with one of those supervised learners.A pandas column with dtype
stringorobjectcontaining strings, or a polars column with dtypeString, is converted to a categorical column. Categorical columns are passed through.If
accept_intis set toTrue, then integer columns are also accepted and converted to categorical. The default value isFalse.Any other type of column is rejected by raising a
RejectColumnexception. Note: theTableVectorizeronly sends string or categorical columns to itslow_cardinality_transformer, regardless of the inputted value ofaccept_int. Therefore it is always safe to use aToCategoricalinstance as thelow_cardinality_transformer.The output of
transformalso always has a Categorical dtype. The categories are not necessarily the same across different calls totransform. Indeed, scikit-learn estimators do not inspect the dtype’s categories but the actual values. Converting to a Categorical is therefore just a way to mark a column and indicate to downstream estimators that this column should be treated as categorical. Ensuring they are encoded consistently, handling unseen categories at test time, etc. is the responsibility of encoders such asOneHotEncoderandLabelEncoder, or of estimators that handle categories themselves such asHistGradientBoostingRegressor.Examples
>>> import pandas as pd >>> from skrub import ToCategorical
A string column is converted to a categorical column.
>>> s = pd.Series(['one', 'two', None], name='c') >>> s 0 one 1 two 2 ... Name: c, dtype: ... >>> to_cat = ToCategorical() >>> to_cat.fit_transform(s) 0 one 1 two 2 ... Name: c, dtype: ... Categories (2, ...): ['one', 'two']
The dtypes (the list of categories) of the outputs of
transformmay vary. This transformer only ensures the dtype is Categorical to mark the column as such for downstream encoders which will perform the actual encoding.>>> s = pd.Series(['four', 'five'], name='c') >>> to_cat.transform(s) 0 four 1 five Name: c, dtype: category Categories (2, ...): ['five', 'four']
Columns that already have a Categorical dtype are passed through:
>>> s = pd.Series(['one', 'two'], name='c', dtype='category') >>> to_cat.fit_transform(s) is s True
Columns that are not strings nor categorical are rejected:
>>> to_cat.fit_transform(pd.Series([1.1, 2.2], name='c')) Traceback (most recent call last): ... skrub.core.RejectColumn: Column 'c' does not contain only strings...
Unless
accept_intis set toTrue, in which case integer columns are accepted:>>> to_cat = ToCategorical(accept_int=True) >>> to_cat.fit_transform(pd.Series([1, 2], name='c')) 0 1 1 2 Name: c, dtype: category Categories (2, int64): [1, 2]
objectcolumns that do not contain only strings are also rejected:>>> s = pd.Series(['one', 1], name='c') >>> to_cat.fit_transform(s) Traceback (most recent call last): ... skrub.core.RejectColumn: Column 'c' does not contain only strings...
No special handling of
StringDtypevsobjectcolumns is done, the behavior is the same aspd.astype('category'): if the input uses the extension dtype, the categories of the output will, too.>>> s = pd.Series(['cat A', 'cat B', None], name='c', dtype='string') >>> s 0 cat A 1 cat B 2 <NA> Name: c, dtype: string >>> to_cat.fit_transform(s) 0 cat A 1 cat B 2 <NA> Name: c, dtype: category Categories (2, string): [...] >>> _.cat.categories.dtype
Polars string columns are converted to the
Categoricaldtype (notEnum). As for pandas, categories may vary across calls totransform.>>> import pytest >>> pl = pytest.importorskip("polars") >>> s = pl.Series('c', ['one', 'two', None]) >>> to_cat.fit_transform(s) shape: (3,) Series: 'c' [cat] [ "one" "two" null ]
Polars Categorical or Enum columns are passed through:
>>> s = pl.Series('c', ['one', 'two'], dtype=pl.Enum(['one', 'two', 'three'])) >>> s shape: (2,) Series: 'c' [enum] [ "one" "two" ] >>> to_cat.fit_transform(s) is s True
Methods
fit(column[, y])Fit the transformer.
fit_transform(column[, y])Fit the encoder and transform a column.
get_feature_names_out([input_features])Get the output feature names.
get_params([deep])Get parameters for this estimator.
set_output(*[, transform])Default no-op implementation for set_output.
set_params(**params)Set the parameters of this estimator.
set_transform_request(*[, column])Configure whether metadata should be requested to be passed to the
transformmethod.transform(column)Transform a column.
- fit(column, y=None, **kwargs)[source]#
Fit the transformer.
This default implementation simply calls
fit_transform()and returnsself.- Parameters:
- columna pandas or polars
Series Unlike most scikit-learn transformers, single-column transformers transform a single column, not a whole dataframe.
- ycolumn or dataframe
Prediction targets.
- **kwargs
Extra named arguments are passed to
self.fit_transform().
- columna pandas or polars
- Returns:
- self
The fitted transformer.
- get_feature_names_out(input_features=None)[source]#
Get the output feature names.
- Parameters:
- input_featuresarray_like of
python:str, default=None Input feature names. Ignored.
- input_featuresarray_like of
- Returns:
python:listofpython:strThe names of the output features.
- get_params(deep=True)[source]#
Get parameters for this estimator.
- Parameters:
- deep
python:bool, default=True If True, will return the parameters for this estimator and contained subobjects that are estimators.
- deep
- Returns:
- params
python:dict Parameter names mapped to their values.
- params
- set_output(*, transform=None)[source]#
Default no-op implementation for set_output.
Skrub transformers already output dataframes of the correct type by default so there is usually no need for set_output to do anything.
Subclasses are of course free to redefine set_output (e.g. by inheriting from
TransformerMixinbefore SingleColumnTransformer).- Parameters:
- transform
python:strorpython:None, default=None Ignored.
- transform
- Returns:
- SingleColumnTransformer
Returns self.
- set_params(**params)[source]#
Set the parameters of this estimator.
The method works on simple estimators as well as on nested objects (such as
Pipeline). The latter have parameters of the form<component>__<parameter>so that it’s possible to update each component of a nested object.- Parameters:
- **params
python:dict Estimator parameters.
- **params
- Returns:
- selfestimator instance
Estimator instance.
- set_transform_request(*, column='$UNCHANGED$')[source]#
Configure whether metadata should be requested to be passed to the
transformmethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed totransformif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it totransform.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.
- Parameters:
- column
python:str,python:True,python:False, orpython:None, default=sklearn.utils.metadata_routing.UNCHANGED Metadata routing for
columnparameter intransform.
- column
- Returns:
- selfobject
The updated object.
Gallery examples#
Encoding: from a dataframe to a numerical matrix for machine learning