训练垃圾邮件识别模型时CountVectorizer报list无lower属性错误
问题原因与解决方案
错误根源
你在预处理阶段把df['Text']和df['Subject']的每个元素转换成了列表类型,但CountVectorizer要求输入的每个样本是字符串。当CountVectorizer执行lower=True时,会尝试对每个样本(也就是列表)调用.lower()方法,而列表没有这个属性,因此触发错误。
具体来看预处理代码中的这部分:
sub = [(my_stemmer_object.stem(i)).lower() for i in sub.split() if i not in stopwords.words('english') and len(i) >= 2] cont = [(my_stemmer_object.stem(i)).lower() for i in cont.split() if i not in stopwords.words('english') and len(i) >= 2] df['Subject'][w] = sub df['Text'][w] = cont
这里生成的sub和cont是词干化后的单词列表,直接赋值给DataFrame的列,导致后续CountVectorizer处理时出错。
修复后的预处理代码
把处理后的单词列表拼接成字符串,同时优化代码效率(避免循环遍历行,改用apply,提前实例化词干器):
import pandas as pd import re from nltk.corpus import stopwords from nltk.stem import SnowballStemmer def getDFcsv(fname): df = pd.read_csv(fname) # 提前实例化词干器,避免循环内重复创建 my_stemmer_object = SnowballStemmer("english") stop_words = set(stopwords.words('english')) # 转成集合,提升查询速度 def process_text(text): text = str(text) # 移除标点、标签、数字和特殊字符 text = re.sub('[^a-zA-Z]', ' ', text) text = re.sub("</?.*?>", " <> ", text) text = re.sub("(\\d|\\W)+", " ", text) # 分词、去停用词、词干化、转小写 words = [ my_stemmer_object.stem(word).lower() for word in text.split() if word not in stop_words and len(word) >= 2 ] # 把列表拼接成字符串 return ' '.join(words) # 用apply批量处理列,比循环高效且避免链式索引警告 df['Subject'] = df['Subject'].apply(process_text) df['Text'] = df['Text'].apply(process_text) return df
关键修复点
- 将词干化后的单词列表用
' '.join(words)拼接成字符串,符合CountVectorizer的输入要求 - 用
apply替代循环遍历行,提升代码效率,同时避免df['Text'][w]这种链式索引的警告 - 提前将停用词转成集合,提升查询速度(集合的
in操作比列表快得多) - 把重复的文本处理逻辑封装成函数,避免代码冗余
补充优化建议
你的createBagOfWords代码可以保持不变,因为现在df['Text']的每个元素都是字符串,CountVectorizer可以正常处理。另外,你预处理时已经做了去停用词、词干化和转小写,可以考虑把CountVectorizer的重复参数去掉,避免冗余操作:
count_vector = CountVectorizer(ngram_range=(1, 1)) # 移除lowercase和stop_words参数
内容的提问来源于stack exchange,提问作者user22344968
相关产品推荐
相关产品推荐

