NLP堆叠分类遇AttributeError报错及格式问题求助
这个问题的核心是ColumnSelector返回的数据格式和TfidfVectorizer期望的输入不匹配,我们一步步拆解和解决:
问题根源分析
当你使用ColumnSelector(cols=(1,))(注意是元组格式)时,它返回的是二维numpy数组(形状为(n_samples, 1))。而TfidfVectorizer需要的输入是一维的字符串序列——比如每个元素都是单独的字符串,而不是包含单个字符串的数组。
所以当TfidfVectorizer尝试对每个输入元素调用.lower()方法时,它拿到的是numpy数组而非字符串,自然就抛出了AttributeError: 'numpy.ndarray' object has no attribute 'lower'。后续把lowercase=False只是绕开了lower方法的问题,但分词器需要匹配字符串,数组依然不符合要求,所以又出现了类型错误。
另外还有一个隐藏问题:你的测试集X_test只有1列,但pipe_1和pipe_2分别选择第1、2列(索引从0开始,对应原X_train的'Mid_level'和'Low_level'),预测时会因为找不到对应列报错,这个也要一起修正。
解决方案
有两种简单的方式修复数据格式问题:
方法1:修改ColumnSelector的cols参数为单个整数
把cols=(1,)改成cols=1(去掉元组),这样ColumnSelector会返回一维的序列(pandas Series或者一维numpy数组),正好符合TfidfVectorizer的要求:
import pandas as pd from sklearn.metrics import accuracy_score from sklearn.pipeline import make_pipeline from mlxtend.feature_selection import ColumnSelector from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.linear_model import LogisticRegression from mlxtend.classifier import StackingClassifier # 创建训练集 start = [ ['apple this is painful Two wrongs make a right ok', 'just a batch of suspicious words and banana', 'another batch of fake words and another apple'], ['Fortune favors the italic sunny sunshine', 'name of a company and then its description', 'is it all sunshine or doomed to fail to no sunshine'], ['this was it when in rome do as the romans and make fortune', 'well again the same thing and those descriptions', 'lets make that work and bring the fortune'], ['Ok this is the last one and then its the end', 'is it the beggining of the end or the end of the beggining', 'allelouia'] ] X_train = pd.DataFrame(start, columns=['High_level', 'Mid_level', 'Low_level']) y_train = ['A', 'B', 'C', 'D'] # 修正测试集:保持和训练集相同的列结构,只填充对应列的数据 X_test = pd.DataFrame( [ ['', 'mostly apple', ''], ['', 'bunch of apple', ''], ['', '', 'lot of fortune'], ['', '', 'make fortune and bring the'], ['', 'beggining of the end', ''] ], columns=['High_level', 'Mid_level', 'Low_level'] ) y_true = ['A', 'A', 'C', 'C', 'D'] # 修改ColumnSelector的cols为单个整数 pipe_1 = make_pipeline(ColumnSelector(cols=1), TfidfVectorizer(min_df=1), LogisticRegression(multi_class='multinomial')) pipe_2 = make_pipeline(ColumnSelector(cols=2), TfidfVectorizer(min_df=1), LogisticRegression(multi_class='multinomial')) sclf = StackingClassifier( classifiers=[pipe_1, pipe_2], meta_classifier=LogisticRegression( solver='lbfgs', multi_class='multinomial', C=1.0, class_weight='balanced', tol=1e-6, max_iter=1000, n_jobs=-1 ) ) # 训练并预测 predictions = sclf.fit(X_train, y_train).predict(X_test) print("预测结果:", predictions) print("准确率:", accuracy_score(y_true, predictions))
方法2:添加FunctionTransformer将二维数组转成一维
如果你需要保留元组形式的cols参数(比如后续要选多列合并处理),可以在管道中添加FunctionTransformer来把二维数组ravel成一维:
from sklearn.preprocessing import FunctionTransformer # 定义转换函数 def flatten_array(X): return X.ravel() pipe_1 = make_pipeline( ColumnSelector(cols=(1,)), FunctionTransformer(flatten_array), TfidfVectorizer(min_df=1), LogisticRegression(multi_class='multinomial') ) pipe_2 = make_pipeline( ColumnSelector(cols=(2,)), FunctionTransformer(flatten_array), TfidfVectorizer(min_df=1), LogisticRegression(multi_class='multinomial') ) # 后续训练和预测代码和方法1一致
为什么之前的ravel()/tolist()无效?
你单独调整数据格式无效是因为这些操作没有整合到管道中——StackingClassifier在训练时会自动把X_train传给每个管道的第一个步骤,所以必须把格式转换逻辑放到管道内部,才能保证数据在每个步骤中都是正确的格式。
内容的提问来源于stack exchange,提问作者Pat

