You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SKLearn中StackingClassifier结合TfidfVectorizer触发AttributeError问题求助

解决StackingClassifier结合NLP与量化模型时的AttributeError问题

在使用StackingClassifier结合NLP模型与量化特征模型时,执行stacking_model.fit(x_train, y_train)触发错误:

AttributeError: 'numpy.ndarray' object has no attribute 'lower'

错误原因

NLP管道中的ColumnTransformer处理单列文本特征后,输出的是二维numpy数组(形状为(n_samples, 1)),而TfidfVectorizer默认期望输入是一维文本序列(如(n_samples,)的字符串数组或Series)。当二维数组传入TfidfVectorizer时,它会将每个子数组视为单个样本,并尝试调用字符串的lower()方法,但numpy数组没有该属性,因此报错。

解决方案

修改NLP管道,添加一个步骤将二维数组转换为一维序列,或者直接提取一维文本列。以下是两种可行的修改方式:

方法1:用FunctionTransformer直接提取一维文本列

替换NLP管道中的ColumnTransformer,直接从DataFrame提取目标文本列并转为一维数组:

from sklearn.preprocessing import FunctionTransformer

# 定义提取文本列的函数
def extract_text(X):
    return X['processed_text'].values

pipe_nlp = Pipeline([
    ('select', FunctionTransformer(extract_text, validate=False)),
    ('tfidf', tfidf_vectorizer),
    ('clf', nlp_model)
])

方法2:在ColumnTransformer后添加扁平化步骤

保留ColumnTransformer,新增一个步骤将二维数组压平为一维:

from sklearn.preprocessing import FunctionTransformer

pipe_nlp = Pipeline([
    ('select', ColumnTransformer([('sel', 'passthrough', NLP_COLS)])),
    # 将(n_samples,1)的二维数组转为(n_samples,)的一维数组
    ('flatten', FunctionTransformer(lambda x: x.ravel(), validate=False)),
    ('tfidf', tfidf_vectorizer),
    ('clf', nlp_model)
])

优化建议

另外,train_test_split可以一次性拆分特征和标签,避免两次拆分可能带来的潜在问题:

x_train, x_test, y_train, y_test = train_test_split(
    combined_tweet_df[['processed_text', 'retweet_count', 'favorite_count', 'num_hashtags', 'num_urls']],
    combined_tweet_df['type'],
    test_size=0.25,
    random_state=4
)

内容的提问来源于stack exchange,提问作者ajaybhatta49

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 15:22:47