使用ColumnTransformer与Pipeline处理玩具数据集时遇ValueError:指定列不在数据框中问题求助
Fixing "ValueError: A given column is not a column of the dataframe" in ColumnTransformer + Pipeline
Let's break down what's causing this error and how to fix it quickly.
The Root Cause
Looking at your code, here's the key issue:
- In step 1, you explicitly drop the
Illnesscolumn from your features dataframe:toy_drop=toy.drop(['Number','Illness'],axis=1) - But in your
ColumnTransformer, you're trying to applyOneHotEncoderto theIllnesscolumn:('ohe',ohe,['City','Gender','Illness']),
Since toy_drop (and thus x_train) no longer contains the Illness column, the transformer throws an error when it can't locate the specified column.
Additional Cleanup Tips
I also noticed a couple of small tweaks to make your code cleaner:
- Your target variable is
Illness, so it should never be included in your feature preprocessing pipeline. - Converting
toy_targetto a dataframe with.to_frame()is unnecessary—train_test_splitworks perfectly with a pandas Series directly.
Corrected Full Code
Here's the fixed version of your workflow:
Step 1: Data Preparation (Cleaned)
import pandas as pd from sklearn.preprocessing import RobustScaler, MinMaxScaler, OneHotEncoder, LabelEncoder, OrdinalEncoder, KBinsDiscretizer from sklearn.compose import ColumnTransformer from sklearn.pipeline import Pipeline from sklearn.linear_model import LinearRegression from sklearn.model_selection import train_test_split toy = pd.read_csv('toy_dataset.csv') # Features: drop irrelevant Number column and target Illness column toy_features = toy.drop(['Number', 'Illness'], axis=1) # Target: keep as a simple Series (no need for to_frame()) toy_target = toy['Illness']
Step 2: Preprocessing Tools (Unchanged)
rb = RobustScaler() normalization = MinMaxScaler() ohe = OneHotEncoder(sparse=False) le = LabelEncoder() oe = OrdinalEncoder() bins = KBinsDiscretizer(n_bins=5, encode='onehot-dense', strategy='uniform')
Step 3: ColumnTransformer & Pipeline (Fixed)
ct_features = ColumnTransformer([ ('normalization', normalization, ['Income']), # Remove Illness from OHE columns—it's our target, not a feature ('ohe', ohe, ['City', 'Gender']), ('bins', bins, ['Age']), ], remainder='drop') pip = Pipeline([ ("ct", ct_features), ('lr', LinearRegression()) ])
Step 4: Train-Test Split & Model Training
x_train, x_test, y_train, y_test = train_test_split(toy_features, toy_target, test_size=0.2, random_state=2021) pip.fit(x_train, y_train)
Quick Side Note
Wait a second—if Illness is a categorical variable (which it sounds like), using LinearRegression might not be the right choice. You probably want to use LogisticRegression for a classification task instead. Just a heads-up to make sure your model matches your problem type!
内容的提问来源于stack exchange,提问作者helloworld
相关产品推荐
相关产品推荐

