You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用make_column_transformer预处理时OneHotEncoder报ValueError: Input contains NaN错误的原因及代码问题咨询

Why OneHotEncoder Throws "ValueError: Input contains NaN" Even After Imputation & Fixes for Your Pipeline

Let's break down why you're hitting this NaN error and fix those pipeline issues step by step.

Why the NaN Error Happens

The core misunderstanding here is how make_column_transformer works: it runs each transformer independently on the original input data, not sequentially.

Your OneHotEncoder is looking directly at the raw Embarked and Sex columns from X_train (which still contain NaNs) instead of the imputed versions from your SimpleImputer step. The imputed data from the first transformer doesn't get passed to the OneHotEncoder—they're separate, parallel processes.

Issues with Your First Pipeline Attempt

  • Parallel processing mismatch: As noted above, each transformer in the column transformer operates on the original columns, not the output of previous transformers. Your imputation only creates imputed copies of the columns, but the OneHotEncoder is still grabbing the raw, NaN-containing columns.
  • Redundant column processing: You're processing columns like Pclass multiple times (once in the SimpleImputer group, once in OrdinalEncoder) which will lead to duplicate features in your final output, causing downstream problems for your model.

Issues with Your Second Pipeline Attempt

  • Losing column names in Pipeline steps: When you run a ColumnTransformer in the first Pipeline step, it outputs a numpy array (not a DataFrame) which has no column names. The second ColumnTransformer tries to reference columns by name (one_hot_col_names), which don't exist in the array—this would throw another error even if you fixed the NaN issue.
  • Pipeline vs ColumnTransformer confusion: Pipelines chain transformations end-to-end (the entire dataset goes through step 1, then step 2), while ColumnTransformers apply different transformations to different columns in parallel. Stacking ColumnTransformers in a Pipeline like this doesn't work because the second one can't map column names to the array input.

The Correct Approach: Per-Column Pipeline Chains

You need to create separate preprocessing pipelines for each column type, then use ColumnTransformer to apply each pipeline to its target columns. This way, each column goes through its own imputation → encoding/scaling steps sequentially.

Here's the fixed code:

from sklearn.metrics import classification_report, confusion_matrix
from sklearn.utils.multiclass import unique_labels
import plotly.figure_factory as ff
import pandas as pd
from sklearn.preprocessing import LabelEncoder, StandardScaler
from sklearn.impute import SimpleImputer
import numpy as np
from sklearn.impute import KNNImputer
from sklearn.model_selection import train_test_split
from sklearn.svm import SVC
from sklearn.tree import DecisionTreeClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OrdinalEncoder, OneHotEncoder
from sklearn.compose import make_column_transformer

random_state = 27912
df_train = pd.read_csv("...")
df_test = pd.read_csv("...")

X_train, X_test, y_train, y_test = train_test_split(
    df_train.drop(["Survived", "Ticket", "Cabin", "Name", "PassengerId"], axis=1),
    df_train["Survived"],
    test_size=0.2,
    random_state=42
)

# Define column groups
numeric_col_names = ["Age", "SibSp", "Parch", "Fare"]
ordinal_col_names = ["Pclass"]
one_hot_col_names = ["Embarked", "Sex"]

# Create separate pipelines for each column type
numeric_pipeline = Pipeline([
    ("imputer", SimpleImputer(strategy="median")),
    ("scaler", StandardScaler())
])

ordinal_pipeline = Pipeline([
    ("imputer", SimpleImputer(strategy="most_frequent")),
    ("encoder", OrdinalEncoder())
])

one_hot_pipeline = Pipeline([
    ("imputer", SimpleImputer(strategy="most_frequent")),
    ("encoder", OneHotEncoder(drop="first", sparse_output=False))  # Drop first to avoid multicollinearity
])

# Combine pipelines with ColumnTransformer
preprocessing_ct = make_column_transformer(
    (numeric_pipeline, numeric_col_names),
    (ordinal_pipeline, ordinal_col_names),
    (one_hot_pipeline, one_hot_col_names)
)

# Fit and transform the training data
X_train_processed = preprocessing_ct.fit_transform(X_train)
# Transform test data (use transform only, no fit!)
X_test_processed = preprocessing_ct.transform(X_test)

Key Fixes Explained:

  1. Per-column sequential processing: Each column group goes through its imputation step first, then the encoding/scaling step—no more NaNs reaching the OneHotEncoder.
  2. No redundant columns: Each column is processed exactly once by its dedicated pipeline.
  3. Clear separation of concerns: Each pipeline handles one type of preprocessing, making the code easier to debug and modify.
  4. drop="first" in OneHotEncoder: Optional but recommended to remove redundant columns (avoids multicollinearity in models like linear regression or SVM).

Once your preprocessing is working, you can chain this with a model in a final pipeline for end-to-end training:

final_pipeline = Pipeline([
    ("preprocessing", preprocessing_ct),
    ("classifier", SVC(random_state=random_state))
])

final_pipeline.fit(X_train, y_train)
y_pred = final_pipeline.predict(X_test)
print(classification_report(y_test, y_pred))

内容的提问来源于stack exchange,提问作者Norhther

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 22:12:44