添加特征后Scikit-learn模型拟合报错:样本数不一致的解决方法
Hey there! Let's break down why you're hitting this error and how to fix it smoothly.
The Root Cause
Your original pipeline works with a single text column (variation, a pandas Series) because your predictors() custom transformer is likely designed to handle one-dimensional text input. When you switch to a 2-column DataFrame (['variation','verified_reviews']), the transformer doesn't know how to process multiple text columns properly. This leads to a misaligned output shape—Scikit-learn expects (n_samples, n_features) but gets something like (2, n_samples) instead, hence the ValueError.
Two Reliable Solutions
Option 1: Process Each Text Column Separately & Combine Features
Use ColumnTransformer to apply your text preprocessing pipeline to each text column individually, then merge the resulting features. This keeps each feature's context intact.
import pandas as pd from sklearn.model_selection import train_test_split from sklearn.pipeline import Pipeline from sklearn.compose import ColumnTransformer # Assume your existing components are defined: predictors(), bow_vector, classifier # First, define a sub-pipeline for text processing text_pipeline = Pipeline([ ('cleaner', predictors()), ('vectorizer', bow_vector) ]) # Use ColumnTransformer to apply this pipeline to both text columns preprocessor = ColumnTransformer( transformers=[ ('text_processing', text_pipeline, ['variation', 'verified_reviews']) ]) # Build the full pipeline full_pipeline = Pipeline([ ('preprocessor', preprocessor), ('classifier', classifier) ]) # Load data and fit df_amazon = pd.read_csv("datasets/amazon_alexa.tsv", sep="\t") X = df_amazon[['variation','verified_reviews']] ylabels = df_amazon['feedback'] X_train, X_test, y_train, y_test = train_test_split(X, ylabels, test_size=0.3) full_pipeline.fit(X_train, y_train)
Option 2: Merge Text Columns Into One
If you don't need to keep the features separate, combine the two text columns into a single string column first. This lets you reuse your original pipeline without modifications.
import pandas as pd from sklearn.model_selection import train_test_split from sklearn.pipeline import Pipeline # Load data and combine text columns df_amazon = pd.read_csv("datasets/amazon_alexa.tsv", sep="\t") df_amazon['combined_text'] = df_amazon['variation'].astype(str) + " " + df_amazon['verified_reviews'].astype(str) # Use the combined column with your original pipeline X = df_amazon['combined_text'] ylabels = df_amazon['feedback'] X_train, X_test, y_train, y_test = train_test_split(X, ylabels, test_size=0.3) pipe = Pipeline([('cleaner', predictors()), ('vectorizer', bow_vector), ('classifier', classifier)]) pipe.fit(X_train, y_train)
Quick Check for Your Custom Transformer
If you want to tweak predictors() to handle DataFrames directly, make sure its fit() and transform() methods accept 2D input (DataFrames) and return a 2D array/matrix with shape (n_samples, processed_features). For example, if it cleans text, loop through each column in the DataFrame and apply cleaning to every row.
内容的提问来源于stack exchange,提问作者Marcel

