You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Sklearn正确编码分类特征?决策树训练遇内存错误

Fixing Memory Errors When Encoding Categorical Features for Sklearn Decision Trees

Hey there, let's break down why you're hitting that memory error and walk through practical fixes—this is a super common pain point when working with large datasets and high-cardinality categorical features!

Why You're Seeing the Memory Error

Your dataset has ~247k samples, and the biggest culprit here is the name feature (along with potentially model or brand). These are high-cardinality features—meaning they have thousands (or even tens of thousands) of unique values.

When you use OneHotEncoder on these features, it creates a new column for every unique value. For example, if name has 50k unique entries, that's 50k new columns added to your dataset. Multiply that by 247k samples, and you're looking at a matrix with billions of entries—way too big for your machine's memory to handle. To make it worse, Sklearn's DecisionTreeClassifier doesn't support sparse matrices, so even if you output a sparse matrix from the encoder, it'll get converted to a dense one under the hood, crashing your process.

Practical Solutions to Fix This

1. Ditch OneHotEncoder for High-Cardinality Features

Instead of one-hot encoding, use a target-based encoder that maps each category to a single numerical value (reducing dimensionality drastically). Sklearn 1.2+ includes TargetEncoder which works great for classification tasks:

from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, TargetEncoder

# Split categorical features by cardinality
low_card_cats = ['gearbox', 'fuelType', 'notRepairedDamage', 'vehicleType']  # Few unique values
high_card_cats = ['name', 'model', 'brand']  # Many unique values
numeric_features = ["yearOfRegistration", "powerPS", "kilometer"]

# Build a preprocessor that handles each feature type appropriately
preprocessor = ColumnTransformer(
    transformers=[
        # One-hot encode low-cardinality features (safe, minimal new columns)
        ('low_cat', OneHotEncoder(handle_unknown='ignore'), low_card_cats),
        # Target encode high-cardinality features (1 column per feature)
        ('high_cat', TargetEncoder(target_type='classification'), high_card_cats),
        # Pass through numeric features as-is
        ('num', 'passthrough', numeric_features)
    ])

# Fit and transform your features (pass the target to TargetEncoder)
X = preprocessor.fit_transform(data.drop('price', axis=1), data['price'])
y = data['price']

2. Remove Unnecessary High-Cardinality Features

The name feature likely adds little unique value beyond what model and brand already provide. Dropping it can immediately cut down on encoding complexity:

# Drop the name feature to reduce cardinality
data = data.drop('name', axis=1)

# Update your categorical feature list
categorical = ['vehicleType', 'gearbox', 'model', 'fuelType', 'brand', 'notRepairedDamage']

3. Optimize Your Dataset's Memory Footprint First

Before encoding, shrink your raw dataset's memory usage to free up space for processing:

# Convert numeric columns to smaller data types (if values fit)
numeric_cols = ['price', 'yearOfRegistration', 'powerPS', 'kilometer']
data[numeric_cols] = data[numeric_cols].astype('int32')

# Convert categorical columns to pandas' category type (reduces object memory overhead)
cat_cols = ['name', 'vehicleType', 'gearbox', 'model', 'fuelType', 'brand', 'notRepairedDamage']
data[cat_cols] = data[cat_cols].astype('category')

This will slash your dataset's memory usage from ~22.6MB to a fraction of that.

4. Avoid DataFrameMapper for Complex Preprocessing

While DataFrameMapper works, Sklearn's native ColumnTransformer is more efficient for combining multiple preprocessing steps, and it integrates seamlessly with the rest of the Sklearn pipeline (like train-test splits and model training).

Final Notes

Once you've preprocessed your data with these fixes, you can safely split into train/test sets and train your decision tree without hitting memory limits:

from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = DecisionTreeClassifier()
model.fit(X_train, y_train)

内容的提问来源于stack exchange,提问作者mandiatutti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 08:47:44