使用make_column_transformer预处理时OneHotEncoder报ValueError: Input contains NaN错误的原因及代码问题咨询
OneHotEncoder Throws "ValueError: Input contains NaN" Even After Imputation & Fixes for Your Pipeline Let's break down why you're hitting this NaN error and fix those pipeline issues step by step.
Why the NaN Error Happens
The core misunderstanding here is how make_column_transformer works: it runs each transformer independently on the original input data, not sequentially.
Your OneHotEncoder is looking directly at the raw Embarked and Sex columns from X_train (which still contain NaNs) instead of the imputed versions from your SimpleImputer step. The imputed data from the first transformer doesn't get passed to the OneHotEncoder—they're separate, parallel processes.
Issues with Your First Pipeline Attempt
- Parallel processing mismatch: As noted above, each transformer in the column transformer operates on the original columns, not the output of previous transformers. Your imputation only creates imputed copies of the columns, but the OneHotEncoder is still grabbing the raw, NaN-containing columns.
- Redundant column processing: You're processing columns like
Pclassmultiple times (once in the SimpleImputer group, once in OrdinalEncoder) which will lead to duplicate features in your final output, causing downstream problems for your model.
Issues with Your Second Pipeline Attempt
- Losing column names in Pipeline steps: When you run a
ColumnTransformerin the first Pipeline step, it outputs a numpy array (not a DataFrame) which has no column names. The secondColumnTransformertries to reference columns by name (one_hot_col_names), which don't exist in the array—this would throw another error even if you fixed the NaN issue. - Pipeline vs ColumnTransformer confusion: Pipelines chain transformations end-to-end (the entire dataset goes through step 1, then step 2), while ColumnTransformers apply different transformations to different columns in parallel. Stacking ColumnTransformers in a Pipeline like this doesn't work because the second one can't map column names to the array input.
The Correct Approach: Per-Column Pipeline Chains
You need to create separate preprocessing pipelines for each column type, then use ColumnTransformer to apply each pipeline to its target columns. This way, each column goes through its own imputation → encoding/scaling steps sequentially.
Here's the fixed code:
from sklearn.metrics import classification_report, confusion_matrix from sklearn.utils.multiclass import unique_labels import plotly.figure_factory as ff import pandas as pd from sklearn.preprocessing import LabelEncoder, StandardScaler from sklearn.impute import SimpleImputer import numpy as np from sklearn.impute import KNNImputer from sklearn.model_selection import train_test_split from sklearn.svm import SVC from sklearn.tree import DecisionTreeClassifier from sklearn.pipeline import Pipeline from sklearn.preprocessing import OrdinalEncoder, OneHotEncoder from sklearn.compose import make_column_transformer random_state = 27912 df_train = pd.read_csv("...") df_test = pd.read_csv("...") X_train, X_test, y_train, y_test = train_test_split( df_train.drop(["Survived", "Ticket", "Cabin", "Name", "PassengerId"], axis=1), df_train["Survived"], test_size=0.2, random_state=42 ) # Define column groups numeric_col_names = ["Age", "SibSp", "Parch", "Fare"] ordinal_col_names = ["Pclass"] one_hot_col_names = ["Embarked", "Sex"] # Create separate pipelines for each column type numeric_pipeline = Pipeline([ ("imputer", SimpleImputer(strategy="median")), ("scaler", StandardScaler()) ]) ordinal_pipeline = Pipeline([ ("imputer", SimpleImputer(strategy="most_frequent")), ("encoder", OrdinalEncoder()) ]) one_hot_pipeline = Pipeline([ ("imputer", SimpleImputer(strategy="most_frequent")), ("encoder", OneHotEncoder(drop="first", sparse_output=False)) # Drop first to avoid multicollinearity ]) # Combine pipelines with ColumnTransformer preprocessing_ct = make_column_transformer( (numeric_pipeline, numeric_col_names), (ordinal_pipeline, ordinal_col_names), (one_hot_pipeline, one_hot_col_names) ) # Fit and transform the training data X_train_processed = preprocessing_ct.fit_transform(X_train) # Transform test data (use transform only, no fit!) X_test_processed = preprocessing_ct.transform(X_test)
Key Fixes Explained:
- Per-column sequential processing: Each column group goes through its imputation step first, then the encoding/scaling step—no more NaNs reaching the OneHotEncoder.
- No redundant columns: Each column is processed exactly once by its dedicated pipeline.
- Clear separation of concerns: Each pipeline handles one type of preprocessing, making the code easier to debug and modify.
drop="first"in OneHotEncoder: Optional but recommended to remove redundant columns (avoids multicollinearity in models like linear regression or SVM).
Once your preprocessing is working, you can chain this with a model in a final pipeline for end-to-end training:
final_pipeline = Pipeline([ ("preprocessing", preprocessing_ct), ("classifier", SVC(random_state=random_state)) ]) final_pipeline.fit(X_train, y_train) y_pred = final_pipeline.predict(X_test) print(classification_report(y_test, y_pred))
内容的提问来源于stack exchange,提问作者Norhther

