Python调用clean_text文本预处理函数报list无lower属性错误求助
错误原因
- 入参类型不匹配:你调用
clean_text时传入的sentences_train、sentences_test是句子列表(多句文本组成的数组),但clean_text当前设计为仅处理单条字符串类型的文本,第一行代码text = text.lower()直接对列表调用字符串方法,自然触发'list' object has no attribute 'lower'报错。你之前单独测试函数正常,应该是测试时传入的是单条字符串而非列表。 - 原
clean_text本身存在多处语法、逻辑错误,即使入参类型正确也无法正常运行:- 字符替换行写为
text- re.sub(...),属于笔误,应将减号改为等号text = re.sub(...) - 缩写展开逻辑的缩进错误:
text = "".join(new_text)写在了for循环内部,每追加一个单词就执行一次拼接,且用空字符串拼接会把所有单词连为无间隔的整串,逻辑完全错误 - 连续调用3次分词器属于冗余操作,且最后词形还原时嵌套map会把单词拆分为单个字符处理,不符合分词还原的预期
- 函数仅在
remove_stopwords为True时才有return语句,若remove_stopwords为False,函数会返回None,导致后续报错
- 字符替换行写为
修复方案
第一步:修复clean_text函数
import re import nltk from nltk.corpus import stopwords # 请提前定义好contractions缩写字典,也可以把remove_stopwords改成函数入参更灵活 remove_stopwords = True contractions = {} # 替换为你自己的缩写映射字典 def clean_text(text): '''处理单条字符串文本''' # 增加类型校验,避免异常输入报错 if not isinstance(text, str): return "" # 转小写 text = text.lower() # 展开缩写 text = text.split() new_text = [] for word in text: if word in contractions: new_text.append(contractions[word]) else: new_text.append(word) # 移到循环外用空格拼接,避免单词粘连 text = " ".join(new_text) # 去除特殊字符、链接等无用内容 text = re.sub(r'https?:\/\/\S+', '', text, flags=re.MULTILINE) text = re.sub(r'\<a href', ' ', text) text = re.sub(r'&', '', text) text = re.sub(r'[_"\-;%()|+&=*%.,!?:#$@\[\]/]',' ', text) text = re.sub(r'<br />', ' ', text) text = re.sub(r'\'', ' ', text) # 合并多余空格 text = re.sub(r'\s+', ' ', text).strip() # 去除停用词 if remove_stopwords: text = text.split() stops = set(stopwords.words("english")) text = [w for w in text if w not in stops] # 分词+词形还原,移除冗余的重复分词操作 lemm = nltk.stem.WordNetLemmatizer() text = [lemm.lemmatize(word) for word in text] # sklearn的向量化器默认输入字符串,所以把分词结果拼接回去;如果你的向量化器支持列表输入,直接返回text即可 return " ".join(text)
第二步:调整调用逻辑,批量处理列表中的所有句子
# 遍历列表对每一条句子单独做预处理 sentences_train = [clean_text(sent) for sent in sentences_train] sentences_test = [clean_text(sent) for sent in sentences_test] X_train = vectorizer.transform(sentences_train) x_test = vectorizer.transform(sentences_test) # 模型评估 score = classifier.score(x_test, y_test) print(colored("Model's Performance", "green")) print("Accuracy:", score)
内容的提问来源于stack exchange,提问作者Ramishka
相关产品推荐
相关产品推荐

