You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

添加特征后Scikit-learn模型拟合报错:样本数不一致的解决方法

How to Fix "Inconsistent Numbers of Samples" Error When Adding a Second Text Feature

Hey there! Let's break down why you're hitting this error and how to fix it smoothly.

The Root Cause

Your original pipeline works with a single text column (variation, a pandas Series) because your predictors() custom transformer is likely designed to handle one-dimensional text input. When you switch to a 2-column DataFrame (['variation','verified_reviews']), the transformer doesn't know how to process multiple text columns properly. This leads to a misaligned output shape—Scikit-learn expects (n_samples, n_features) but gets something like (2, n_samples) instead, hence the ValueError.

Two Reliable Solutions

Option 1: Process Each Text Column Separately & Combine Features

Use ColumnTransformer to apply your text preprocessing pipeline to each text column individually, then merge the resulting features. This keeps each feature's context intact.

import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer

# Assume your existing components are defined: predictors(), bow_vector, classifier
# First, define a sub-pipeline for text processing
text_pipeline = Pipeline([
    ('cleaner', predictors()),
    ('vectorizer', bow_vector)
])

# Use ColumnTransformer to apply this pipeline to both text columns
preprocessor = ColumnTransformer(
    transformers=[
        ('text_processing', text_pipeline, ['variation', 'verified_reviews'])
    ])

# Build the full pipeline
full_pipeline = Pipeline([
    ('preprocessor', preprocessor),
    ('classifier', classifier)
])

# Load data and fit
df_amazon = pd.read_csv("datasets/amazon_alexa.tsv", sep="\t")
X = df_amazon[['variation','verified_reviews']]
ylabels = df_amazon['feedback']
X_train, X_test, y_train, y_test = train_test_split(X, ylabels, test_size=0.3)

full_pipeline.fit(X_train, y_train)

Option 2: Merge Text Columns Into One

If you don't need to keep the features separate, combine the two text columns into a single string column first. This lets you reuse your original pipeline without modifications.

import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline

# Load data and combine text columns
df_amazon = pd.read_csv("datasets/amazon_alexa.tsv", sep="\t")
df_amazon['combined_text'] = df_amazon['variation'].astype(str) + " " + df_amazon['verified_reviews'].astype(str)

# Use the combined column with your original pipeline
X = df_amazon['combined_text']
ylabels = df_amazon['feedback']
X_train, X_test, y_train, y_test = train_test_split(X, ylabels, test_size=0.3)

pipe = Pipeline([('cleaner', predictors()), ('vectorizer', bow_vector), ('classifier', classifier)])
pipe.fit(X_train, y_train)

Quick Check for Your Custom Transformer

If you want to tweak predictors() to handle DataFrames directly, make sure its fit() and transform() methods accept 2D input (DataFrames) and return a 2D array/matrix with shape (n_samples, processed_features). For example, if it cleans text, loop through each column in the DataFrame and apply cleaning to every row.

内容的提问来源于stack exchange,提问作者Marcel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:34:03