You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

训练决策树分类器时train_test_split无法正常工作的技术求助

Fixing train_test_split Issues with Mixed Sparse & Dense Features for Decision Trees

Hey there! It sounds like your problem boils down to mixing sparse matrices (from your OneHotEncoder) and regular numpy arrays—train_test_split can't handle that mismatch directly. Let's break down how to fix this step by step.

First, Fix Your OHL Function

Your code snippet cuts off, so let's make sure the OneHotEncoder wrapper is complete and correct (note the variable name typo you had with Labe...):

from sklearn.preprocessing import LabelEncoder, OneHotEncoder
import scipy.sparse

def OHL(x, column):
    # OneHotEncoder helper for categorical columns
    le = LabelEncoder()
    # Convert to string to handle any non-numeric categorical values
    labeled = le.fit_transform(x[column].astype(str))
    # Reshape to 2D array—OneHotEncoder expects 2D input
    labeled = labeled.reshape(-1, 1)
    enc = OneHotEncoder(sparse_output=True)  # Keep sparse output to save memory
    return enc.fit_transform(labeled)

The Core Problem: Merging Sparse & Dense Features

train_test_split needs a single, consistent input for your features. You can't just combine sparse matrices and numpy arrays with regular concatenation—use scipy.sparse.hstack to merge them into one big sparse matrix instead.

Example Workflow

Let's assume your dataset is in a pandas DataFrame called df:

from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
import pandas as pd

# Step 1: Process categorical features into sparse matrices
sparse_EN = OHL(df, 'EN')
sparse_SN = OHL(df, 'SN')

# Step 2: Extract dense numerical features, convert to sparse matrix
dense_features = df[['JT', 'FT', 'PW', 'YR', 'LO', 'LA']].values
sparse_dense = scipy.sparse.csr_matrix(dense_features)

# Step 3: Merge all features into one sparse matrix
X = scipy.sparse.hstack([sparse_EN, sparse_SN, sparse_dense])

# Step 4: Grab your target variable
y = df['CS'].values

# Now train_test_split works perfectly!
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

# Train your decision tree
clf = DecisionTreeClassifier()
clf.fit(X_train, y_train)

A Cleaner Alternative: Use ColumnTransformer

Instead of manually handling sparse/dense merges, let scikit-learn do the work with ColumnTransformer—it's designed exactly for mixed feature types:

from sklearn.compose import ColumnTransformer

# Define preprocessing pipeline: handle categoricals and numerics separately
preprocessor = ColumnTransformer(
    transformers=[
        # Apply OneHotEncoder to categorical columns
        ('categorical', OneHotEncoder(sparse_output=True), ['EN', 'SN']),
        # Pass through numerical columns without changes
        ('numerical', 'passthrough', ['JT', 'FT', 'PW', 'YR', 'LO', 'LA'])
    ]
)

# Transform your entire DataFrame in one step
X = preprocessor.fit_transform(df)
y = df['CS'].values

# Split and train as before
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
clf = DecisionTreeClassifier()
clf.fit(X_train, y_train)

This approach is more maintainable and avoids manual errors from handling sparse matrices directly.

内容的提问来源于stack exchange,提问作者Ekkasit Smithipanon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:43:44