Python文本清洗与高频词统计问题:停用词未被过滤求助
问题分析与解决方法
嘿,刚入门Python遇到这种问题太正常啦,我来帮你捋清楚问题出在哪,一步步解决~
为什么停用词还会出现在结果里?
你遇到的问题主要是两个细节没处理到位:
- 停用词列表没统一转小写:你的代码里把待过滤的单词转成了
lower(),但停用词列表里的词还是原格式(比如停用词文件里如果有The或者混合大小写的词,就匹配不上小写的the)。 - 文本标点没提前处理:用
split()分割文本时,像the,、High Representative’这种带标点的词会被当成一个整体,转小写后是the,,和停用词里的the不匹配,所以没被过滤;之后用re.findall(r'\w+', cleantxt)提取单词时,又把the从the,里拆出来,导致统计时出现停用词。
修复步骤(一步步来)
1. 先处理停用词列表
把停用词全部转成小写,确保和待过滤单词的格式完全一致:
with open("stopwords.txt") as f: stopwords = f.readlines() # 先去掉换行符,再统一转小写 stopwords = [x.strip().lower() for x in stopwords]
2. 提前提取文本里的纯单词(去掉标点)
不要直接用split(),而是先用正则提取所有字母数字组成的单词,这样带标点的词会被拆成纯单词,方便后续过滤:
import re from collections import Counter # 读取待清洗文本 with open("text_to_be_cleaned.txt", "r", encoding="utf-8") as f: txt = f.read() # 提取所有纯单词,同时转成小写 words = re.findall(r'\w+', txt.lower())
3. 过滤停用词
现在直接过滤就可以,因为所有单词和停用词都是小写,匹配更准确:
clean_words = [word for word in words if word not in stopwords]
4. 统计频率最高的25个词
这一步逻辑和你原来的差不多,但现在已经是纯小写的干净单词了,直接统计就行:
word_counts = Counter(clean_words).most_common(25) print(word_counts)
完整可运行代码
把上面的步骤整合起来,就是完整的修复后代码:
#!/usr/bin/python # -*- coding: utf-8 -*- import re from collections import Counter # 读取并处理停用词列表 with open("stopwords.txt", "r", encoding="utf-8") as f: # 去重+转小写+去换行,提升过滤效率 stopwords = list(set([line.strip().lower() for line in f.readlines()])) # 读取待清洗文本并提取纯单词(统一转小写) with open("text_to_be_cleaned.txt", "r", encoding="utf-8") as f: txt = f.read() words = re.findall(r'\w+', txt.lower()) # 过滤停用词 clean_words = [word for word in words if word not in stopwords] # 统计Top25高频词并打印 top25_words = Counter(clean_words).most_common(25) print("Top 25高频词:") for word, count in top25_words: print(f"{word}: {count}")
额外小提示
- 记得指定文件编码为
utf-8,避免读取带特殊字符的文本时出现乱码(你原来的iso-8859-15可能不适合处理示例里的特殊引号这类字符)。 - 给停用词列表转成集合去重,可以提升过滤的效率哦~
内容的提问来源于stack exchange,提问作者Cold2Breath
相关产品推荐
相关产品推荐

