You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

训练垃圾邮件识别模型时CountVectorizer报list无lower属性错误

问题原因与解决方案

错误根源

你在预处理阶段把df['Text']和df['Subject']的每个元素转换成了列表类型,但CountVectorizer要求输入的每个样本是字符串。当CountVectorizer执行lower=True时,会尝试对每个样本(也就是列表)调用.lower()方法,而列表没有这个属性,因此触发错误。

具体来看预处理代码中的这部分:

sub = [(my_stemmer_object.stem(i)).lower() for i in sub.split() if i not in stopwords.words('english') and len(i) >= 2]
cont = [(my_stemmer_object.stem(i)).lower() for i in cont.split() if i not in stopwords.words('english') and len(i) >= 2]
df['Subject'][w] = sub
df['Text'][w] = cont

这里生成的sub和cont是词干化后的单词列表,直接赋值给DataFrame的列,导致后续CountVectorizer处理时出错。

修复后的预处理代码

把处理后的单词列表拼接成字符串,同时优化代码效率(避免循环遍历行,改用apply,提前实例化词干器):

import pandas as pd
import re
from nltk.corpus import stopwords
from nltk.stem import SnowballStemmer

def getDFcsv(fname):
    df = pd.read_csv(fname)
    # 提前实例化词干器,避免循环内重复创建
    my_stemmer_object = SnowballStemmer("english")
    stop_words = set(stopwords.words('english'))  # 转成集合,提升查询速度

    def process_text(text):
        text = str(text)
        # 移除标点、标签、数字和特殊字符
        text = re.sub('[^a-zA-Z]', ' ', text)
        text = re.sub("</?.*?>", " <> ", text)
        text = re.sub("(\\d|\\W)+", " ", text)
        # 分词、去停用词、词干化、转小写
        words = [
            my_stemmer_object.stem(word).lower() 
            for word in text.split() 
            if word not in stop_words and len(word) >= 2
        ]
        # 把列表拼接成字符串
        return ' '.join(words)

    # 用apply批量处理列,比循环高效且避免链式索引警告
    df['Subject'] = df['Subject'].apply(process_text)
    df['Text'] = df['Text'].apply(process_text)

    return df

关键修复点

  • 将词干化后的单词列表用' '.join(words)拼接成字符串,符合CountVectorizer的输入要求
  • 用apply替代循环遍历行,提升代码效率,同时避免df['Text'][w]这种链式索引的警告
  • 提前将停用词转成集合,提升查询速度(集合的in操作比列表快得多)
  • 把重复的文本处理逻辑封装成函数,避免代码冗余

补充优化建议

你的createBagOfWords代码可以保持不变,因为现在df['Text']的每个元素都是字符串,CountVectorizer可以正常处理。另外,你预处理时已经做了去停用词、词干化和转小写,可以考虑把CountVectorizer的重复参数去掉,避免冗余操作:

count_vector = CountVectorizer(ngram_range=(1, 1))  # 移除lowercase和stop_words参数

内容的提问来源于stack exchange,提问作者user22344968

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 05:30:14