You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用ColumnTransformer与Pipeline处理玩具数据集时遇ValueError:指定列不在数据框中问题求助

Fixing "ValueError: A given column is not a column of the dataframe" in ColumnTransformer + Pipeline

Let's break down what's causing this error and how to fix it quickly.

The Root Cause

Looking at your code, here's the key issue:

  • In step 1, you explicitly drop the Illness column from your features dataframe:
    toy_drop=toy.drop(['Number','Illness'],axis=1)
    
  • But in your ColumnTransformer, you're trying to apply OneHotEncoder to the Illness column:
    ('ohe',ohe,['City','Gender','Illness']), 
    

Since toy_drop (and thus x_train) no longer contains the Illness column, the transformer throws an error when it can't locate the specified column.

Additional Cleanup Tips

I also noticed a couple of small tweaks to make your code cleaner:

  1. Your target variable is Illness, so it should never be included in your feature preprocessing pipeline.
  2. Converting toy_target to a dataframe with .to_frame() is unnecessary—train_test_split works perfectly with a pandas Series directly.

Corrected Full Code

Here's the fixed version of your workflow:

Step 1: Data Preparation (Cleaned)

import pandas as pd
from sklearn.preprocessing import RobustScaler, MinMaxScaler, OneHotEncoder, LabelEncoder, OrdinalEncoder, KBinsDiscretizer
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split

toy = pd.read_csv('toy_dataset.csv')
# Features: drop irrelevant Number column and target Illness column
toy_features = toy.drop(['Number', 'Illness'], axis=1)
# Target: keep as a simple Series (no need for to_frame())
toy_target = toy['Illness']

Step 2: Preprocessing Tools (Unchanged)

rb = RobustScaler()
normalization = MinMaxScaler()
ohe = OneHotEncoder(sparse=False)
le = LabelEncoder()
oe = OrdinalEncoder()
bins = KBinsDiscretizer(n_bins=5, encode='onehot-dense', strategy='uniform')

Step 3: ColumnTransformer & Pipeline (Fixed)

ct_features = ColumnTransformer([
    ('normalization', normalization, ['Income']), 
    # Remove Illness from OHE columns—it's our target, not a feature
    ('ohe', ohe, ['City', 'Gender']), 
    ('bins', bins, ['Age']), 
], remainder='drop')

pip = Pipeline([ 
    ("ct", ct_features), 
    ('lr', LinearRegression())
])

Step 4: Train-Test Split & Model Training

x_train, x_test, y_train, y_test = train_test_split(toy_features, toy_target, test_size=0.2, random_state=2021)
pip.fit(x_train, y_train)

Quick Side Note

Wait a second—if Illness is a categorical variable (which it sounds like), using LinearRegression might not be the right choice. You probably want to use LogisticRegression for a classification task instead. Just a heads-up to make sure your model matches your problem type!

内容的提问来源于stack exchange,提问作者helloworld

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 02:42:29