训练决策树分类器时train_test_split无法正常工作的技术求助
train_test_split Issues with Mixed Sparse & Dense Features for Decision Trees Hey there! It sounds like your problem boils down to mixing sparse matrices (from your OneHotEncoder) and regular numpy arrays—train_test_split can't handle that mismatch directly. Let's break down how to fix this step by step.
First, Fix Your OHL Function
Your code snippet cuts off, so let's make sure the OneHotEncoder wrapper is complete and correct (note the variable name typo you had with Labe...):
from sklearn.preprocessing import LabelEncoder, OneHotEncoder import scipy.sparse def OHL(x, column): # OneHotEncoder helper for categorical columns le = LabelEncoder() # Convert to string to handle any non-numeric categorical values labeled = le.fit_transform(x[column].astype(str)) # Reshape to 2D array—OneHotEncoder expects 2D input labeled = labeled.reshape(-1, 1) enc = OneHotEncoder(sparse_output=True) # Keep sparse output to save memory return enc.fit_transform(labeled)
The Core Problem: Merging Sparse & Dense Features
train_test_split needs a single, consistent input for your features. You can't just combine sparse matrices and numpy arrays with regular concatenation—use scipy.sparse.hstack to merge them into one big sparse matrix instead.
Example Workflow
Let's assume your dataset is in a pandas DataFrame called df:
from sklearn.model_selection import train_test_split from sklearn.tree import DecisionTreeClassifier import pandas as pd # Step 1: Process categorical features into sparse matrices sparse_EN = OHL(df, 'EN') sparse_SN = OHL(df, 'SN') # Step 2: Extract dense numerical features, convert to sparse matrix dense_features = df[['JT', 'FT', 'PW', 'YR', 'LO', 'LA']].values sparse_dense = scipy.sparse.csr_matrix(dense_features) # Step 3: Merge all features into one sparse matrix X = scipy.sparse.hstack([sparse_EN, sparse_SN, sparse_dense]) # Step 4: Grab your target variable y = df['CS'].values # Now train_test_split works perfectly! X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 ) # Train your decision tree clf = DecisionTreeClassifier() clf.fit(X_train, y_train)
A Cleaner Alternative: Use ColumnTransformer
Instead of manually handling sparse/dense merges, let scikit-learn do the work with ColumnTransformer—it's designed exactly for mixed feature types:
from sklearn.compose import ColumnTransformer # Define preprocessing pipeline: handle categoricals and numerics separately preprocessor = ColumnTransformer( transformers=[ # Apply OneHotEncoder to categorical columns ('categorical', OneHotEncoder(sparse_output=True), ['EN', 'SN']), # Pass through numerical columns without changes ('numerical', 'passthrough', ['JT', 'FT', 'PW', 'YR', 'LO', 'LA']) ] ) # Transform your entire DataFrame in one step X = preprocessor.fit_transform(df) y = df['CS'].values # Split and train as before X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) clf = DecisionTreeClassifier() clf.fit(X_train, y_train)
This approach is more maintainable and avoids manual errors from handling sparse matrices directly.
内容的提问来源于stack exchange,提问作者Ekkasit Smithipanon

