如何为Sklearn正确编码分类特征?决策树训练遇内存错误
Hey there, let's break down why you're hitting that memory error and walk through practical fixes—this is a super common pain point when working with large datasets and high-cardinality categorical features!
Why You're Seeing the Memory Error
Your dataset has ~247k samples, and the biggest culprit here is the name feature (along with potentially model or brand). These are high-cardinality features—meaning they have thousands (or even tens of thousands) of unique values.
When you use OneHotEncoder on these features, it creates a new column for every unique value. For example, if name has 50k unique entries, that's 50k new columns added to your dataset. Multiply that by 247k samples, and you're looking at a matrix with billions of entries—way too big for your machine's memory to handle. To make it worse, Sklearn's DecisionTreeClassifier doesn't support sparse matrices, so even if you output a sparse matrix from the encoder, it'll get converted to a dense one under the hood, crashing your process.
Practical Solutions to Fix This
1. Ditch OneHotEncoder for High-Cardinality Features
Instead of one-hot encoding, use a target-based encoder that maps each category to a single numerical value (reducing dimensionality drastically). Sklearn 1.2+ includes TargetEncoder which works great for classification tasks:
from sklearn.compose import ColumnTransformer from sklearn.preprocessing import OneHotEncoder, TargetEncoder # Split categorical features by cardinality low_card_cats = ['gearbox', 'fuelType', 'notRepairedDamage', 'vehicleType'] # Few unique values high_card_cats = ['name', 'model', 'brand'] # Many unique values numeric_features = ["yearOfRegistration", "powerPS", "kilometer"] # Build a preprocessor that handles each feature type appropriately preprocessor = ColumnTransformer( transformers=[ # One-hot encode low-cardinality features (safe, minimal new columns) ('low_cat', OneHotEncoder(handle_unknown='ignore'), low_card_cats), # Target encode high-cardinality features (1 column per feature) ('high_cat', TargetEncoder(target_type='classification'), high_card_cats), # Pass through numeric features as-is ('num', 'passthrough', numeric_features) ]) # Fit and transform your features (pass the target to TargetEncoder) X = preprocessor.fit_transform(data.drop('price', axis=1), data['price']) y = data['price']
2. Remove Unnecessary High-Cardinality Features
The name feature likely adds little unique value beyond what model and brand already provide. Dropping it can immediately cut down on encoding complexity:
# Drop the name feature to reduce cardinality data = data.drop('name', axis=1) # Update your categorical feature list categorical = ['vehicleType', 'gearbox', 'model', 'fuelType', 'brand', 'notRepairedDamage']
3. Optimize Your Dataset's Memory Footprint First
Before encoding, shrink your raw dataset's memory usage to free up space for processing:
# Convert numeric columns to smaller data types (if values fit) numeric_cols = ['price', 'yearOfRegistration', 'powerPS', 'kilometer'] data[numeric_cols] = data[numeric_cols].astype('int32') # Convert categorical columns to pandas' category type (reduces object memory overhead) cat_cols = ['name', 'vehicleType', 'gearbox', 'model', 'fuelType', 'brand', 'notRepairedDamage'] data[cat_cols] = data[cat_cols].astype('category')
This will slash your dataset's memory usage from ~22.6MB to a fraction of that.
4. Avoid DataFrameMapper for Complex Preprocessing
While DataFrameMapper works, Sklearn's native ColumnTransformer is more efficient for combining multiple preprocessing steps, and it integrates seamlessly with the rest of the Sklearn pipeline (like train-test splits and model training).
Final Notes
Once you've preprocessed your data with these fixes, you can safely split into train/test sets and train your decision tree without hitting memory limits:
from sklearn.model_selection import train_test_split from sklearn.tree import DecisionTreeClassifier X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) model = DecisionTreeClassifier() model.fit(X_train, y_train)
内容的提问来源于stack exchange,提问作者mandiatutti

