You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CountVectorizer与ColumnTransformer联用报错及可扩展解决方案咨询

问题原因

报错的核心是:当你用列表["corpus"]指定文本列时,ColumnTransformer会将该列以二维DataFrame(形状(5,1))的形式传递给CountVectorizer。而CountVectorizer默认会把二维输入的每一列当作一个文档,因此最终输出的特征矩阵形状是(1, n_features),和remainder部分输出的(5,1)矩阵在拼接时维度不匹配(前者行数为1,后者为5),触发ValueError。

解决方案

方案1:直接传递列名字符串(推荐,简洁高效)

将文本列的指定方式从列表改为单个字符串,ColumnTransformer会自动传递一维的Series给CountVectorizer,保证输出矩阵的行数和原数据一致。

from sklearn.compose import ColumnTransformer
from sklearn.feature_extraction.text import CountVectorizer
import pandas as pd

# 构建样本数据
df = pd.DataFrame({
    'corpus': ['This is the first document.', 'This document is the second document.', 'And this is the third one.',
               'Is this the first document?', 'I have the fourth document'],
    'word_length': [27, 37, 26, 27, 26]
})

count_transformer = CountVectorizer()

# 直接传入列名字符串"corpus",而非列表
ct = ColumnTransformer(transformers=[
    ("count", count_transformer, "corpus")],
    remainder='passthrough')

# 执行转换
result = ct.fit_transform(df)
print(result.toarray())

运行后会得到和你手动拼接一致的结果,且扩展性强——后续新增文本列或数值列,只需在transformers中添加对应规则即可。

方案2:用FunctionTransformer适配二维输入(适用于必须传递列列表的场景)

如果因业务需求必须用列表指定列(比如批量处理多列文本),可以通过FunctionTransformer将二维DataFrame转为一维数组,再传递给CountVectorizer:

from sklearn.compose import ColumnTransformer
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.preprocessing import FunctionTransformer
from sklearn.pipeline import Pipeline
import pandas as pd

df = pd.DataFrame({
    'corpus': ['This is the first document.', 'This document is the second document.', 'And this is the third one.',
               'Is this the first document?', 'I have the fourth document'],
    'word_length': [27, 37, 26, 27, 26]
})

# 定义转换函数:将二维DataFrame转为一维文本数组
def flatten_text(X):
    return X.iloc[:, 0].values

# 构建文本处理流水线:先转一维,再向量化
text_pipeline = Pipeline([
    ("flatten", FunctionTransformer(flatten_text)),
    ("vectorize", CountVectorizer())
])

ct = ColumnTransformer(transformers=[
    ("count", text_pipeline, ["corpus"])],
    remainder='passthrough')

result = ct.fit_transform(df)
print(result.toarray())

这个方案可以轻松扩展到多列文本处理——只需调整flatten_text函数,将多列文本合并或分别处理即可。

内容的提问来源于stack exchange,提问作者Chukwudi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 19:55:11