You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python调用clean_text文本预处理函数报list无lower属性错误求助

错误原因
  • 入参类型不匹配:你调用clean_text时传入的sentences_train、sentences_test是句子列表(多句文本组成的数组),但clean_text当前设计为仅处理单条字符串类型的文本,第一行代码text = text.lower()直接对列表调用字符串方法,自然触发'list' object has no attribute 'lower'报错。你之前单独测试函数正常,应该是测试时传入的是单条字符串而非列表。
  • 原clean_text本身存在多处语法、逻辑错误,即使入参类型正确也无法正常运行:
    • 字符替换行写为text- re.sub(...),属于笔误,应将减号改为等号text = re.sub(...)
    • 缩写展开逻辑的缩进错误:text = "".join(new_text)写在了for循环内部,每追加一个单词就执行一次拼接,且用空字符串拼接会把所有单词连为无间隔的整串,逻辑完全错误
    • 连续调用3次分词器属于冗余操作,且最后词形还原时嵌套map会把单词拆分为单个字符处理,不符合分词还原的预期
    • 函数仅在remove_stopwords为True时才有return语句,若remove_stopwords为False,函数会返回None,导致后续报错
修复方案

第一步:修复clean_text函数

import re
import nltk
from nltk.corpus import stopwords
# 请提前定义好contractions缩写字典,也可以把remove_stopwords改成函数入参更灵活
remove_stopwords = True
contractions = {} # 替换为你自己的缩写映射字典

def clean_text(text): 
    '''处理单条字符串文本'''
    # 增加类型校验,避免异常输入报错
    if not isinstance(text, str):
        return ""
    
    # 转小写
    text = text.lower()

    # 展开缩写
    text = text.split()
    new_text = []
    for word in text:
        if word in contractions:
            new_text.append(contractions[word])
        else:
            new_text.append(word)
    # 移到循环外用空格拼接,避免单词粘连
    text = " ".join(new_text)
  
    # 去除特殊字符、链接等无用内容
    text = re.sub(r'https?:\/\/\S+', '', text, flags=re.MULTILINE) 
    text = re.sub(r'\<a href', ' ', text)
    text = re.sub(r'&amp;', '', text)
    text = re.sub(r'[_"\-;%()|+&=*%.,!?:#$@\[\]/]',' ', text)
    text = re.sub(r'<br />', ' ', text)
    text = re.sub(r'\'', ' ', text)
    # 合并多余空格
    text = re.sub(r'\s+', ' ', text).strip()

    # 去除停用词
    if remove_stopwords:
        text = text.split()
        stops = set(stopwords.words("english"))
        text = [w for w in text if w not in stops]
    
    # 分词+词形还原,移除冗余的重复分词操作
    lemm = nltk.stem.WordNetLemmatizer()
    text = [lemm.lemmatize(word) for word in text]
    
    # sklearn的向量化器默认输入字符串,所以把分词结果拼接回去;如果你的向量化器支持列表输入,直接返回text即可
    return " ".join(text)

第二步:调整调用逻辑,批量处理列表中的所有句子

# 遍历列表对每一条句子单独做预处理
sentences_train = [clean_text(sent) for sent in sentences_train]
sentences_test = [clean_text(sent) for sent in sentences_test]

X_train = vectorizer.transform(sentences_train) 
x_test = vectorizer.transform(sentences_test)

# 模型评估
score = classifier.score(x_test, y_test)

print(colored("Model's Performance", "green")) 
print("Accuracy:", score)

内容的提问来源于stack exchange,提问作者Ramishka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 01:24:05