filter_names#
- skrub.selectors.filter_names(predicate, *args, **kwargs)[source]#
Select columns based only on their name.
The predicate takes as input the column name and must return True if it should be selected and False otherwise.
- Parameters:
- predicate
python:callable() A function that takes a column name (string) and optional extra arguments, returning
Trueto select the column orFalseto exclude it. Signature:predicate(col_name, *args, **kwargs) -> bool- *args
python:tuple Extra positional arguments passed to the predicate. Using explicit arguments (instead of closures) helps with pickling the selector.
- **kwargs
python:dict Extra keyword arguments passed to the predicate. Using explicit arguments (instead of closures) helps with pickling the selector.
- predicate
- Returns:
- Selector
A
NameFilterselector that matches columns wherepredicate(name, *args, **kwargs)returnsTrue.
See also
Notes
filter_namespasses the column NAME (a string), whilefilterpasses the column OBJECT (Series/DataFrame)For common patterns, prefer simpler selectors:
s.glob('*_mm')instead ofs.filter_names(lambda n: n.endswith('_mm'))s.regex(r'col_\d+')instead ofs.filter_names(lambda n: re.match(...))
To pickle the selector, use importable functions as predicates rather than lambdas. Pass parameters via
*argsor**kwargs:# Picklable (importable function + explicit args) s.filter_names(str.endswith, '_mm') # Picklable alternative (with functools.partial) import functools s.filter_names(functools.partial(str.endswith, '_mm')) # NOT picklable (closure) suffix = '_mm' s.filter_names(lambda name: name.endswith(suffix))
Examples
>>> from skrub import selectors as s >>> import pandas as pd >>> df = pd.DataFrame( ... { ... "height_mm": [297.0, 420.0], ... "width_mm": [210.0, 297.0], ... "kind": ["A4", "A3"], ... "ID": [4, 3], ... } ... ) >>> df height_mm width_mm kind ID 0 297.0 210.0 A4 4 1 420.0 297.0 A3 3
Prefer using *args and **kwargs to pass extra arguments to the predicate, rather than defining a dynamic function which may cause pickling errors.
Select columns ending with a suffix (using lambda):
>>> selector = s.filter_names(lambda name: name.endswith('_mm')) >>> s.select(df, selector) height_mm width_mm 0 297.0 210.0 1 420.0 297.0
Use an importable function with explicit arguments (picklable):
>>> selector = s.filter_names(str.endswith, '_mm') >>> selector filter_names(str.endswith, '_mm')
>>> s.select(df, selector) height_mm width_mm 0 297.0 210.0 1 420.0 297.0
>>> import pickle >>> _ = pickle.dumps(selector) # Pickling works!
Select columns with uppercase names:
>>> s.select(df, s.filter_names(str.isupper)) ID 0 4 1 3
Combine with type selectors:
>>> s.select(df, s.numeric() & s.filter_names(str.islower)) height_mm width_mm 0 297.0 210.0 1 420.0 297.0