停用词移除失效问题求助:Python代码调试及解决方案咨询
修复停用词未被移除的解决方案
核心问题排查点
- 自定义停用词未正确合并到基础停用词集合
- 文本处理顺序导致停用词匹配失败(比如先过滤停用词再做小写转换/俚语替换)
- 自定义停用词未做与文本一致的预处理(比如没转小写)
具体修复步骤
1. 对齐自定义停用词与文本预处理逻辑
先把自定义停用词列表统一转成小写,和clean_text里的小写转换步骤保持一致:
custom_stopwords = ["dok", "其他自定义停用词"] # 统一转小写 custom_stopwords = [word.lower() for word in custom_stopwords]
2. 正确合并基础与自定义停用词
不要直接替换基础停用词,而是合并两者的集合,避免覆盖原有停用词:
from Sastrawi.StopWordRemover.StopWordRemoverFactory import StopWordRemoverFactory factory = StopWordRemoverFactory() # 获取基础停用词集合 base_stopwords = factory.get_stop_words() # 合并自定义停用词 merged_stopwords = set(base_stopwords + custom_stopwords) # 创建包含自定义词的停用词移除器 stopword_remover = factory.create_stop_word_remover(merged_stopwords)
3. 调整文本处理顺序(关键)
如果停用词过滤步骤在小写转换/俚语替换之前,会导致匹配失败。正确的处理顺序应该是:
- 小写转换
- 俚语替换
- 停用词过滤
- 词干提取
修改后的clean_text示例:
def clean_text(text): # 1. 小写转换 text = text.lower() # 2. 俚语替换(假设你有slang_dict映射字典) for slang, formal in slang_dict.items(): text = text.replace(slang, formal) # 3. 停用词过滤 text = stopword_remover.remove(text) # 4. 词干提取(假设你有已初始化的stemmer实例) words = text.split() stemmed_words = [stemmer.stem(word) for word in words] return ' '.join(stemmed_words)
4. 验证停用词集合
在代码里打印验证,确认目标停用词已被包含:
print("dok" in merged_stopwords) # 正常应该返回True
测试效果
输入:Dok,anak saya sudah imunisasi DPT
修复后输出会移除"dok",得到类似anak saya sudah imunisasi dpt(词干提取后会进一步处理)的结果。
内容的提问来源于stack exchange,提问作者abbym
相关产品推荐
相关产品推荐

