You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

停用词移除失效问题求助:Python代码调试及解决方案咨询

修复停用词未被移除的解决方案

核心问题排查点

  • 自定义停用词未正确合并到基础停用词集合
  • 文本处理顺序导致停用词匹配失败(比如先过滤停用词再做小写转换/俚语替换)
  • 自定义停用词未做与文本一致的预处理(比如没转小写)

具体修复步骤

1. 对齐自定义停用词与文本预处理逻辑

先把自定义停用词列表统一转成小写,和clean_text里的小写转换步骤保持一致:

custom_stopwords = ["dok", "其他自定义停用词"]
# 统一转小写
custom_stopwords = [word.lower() for word in custom_stopwords]

2. 正确合并基础与自定义停用词

不要直接替换基础停用词,而是合并两者的集合,避免覆盖原有停用词:

from Sastrawi.StopWordRemover.StopWordRemoverFactory import StopWordRemoverFactory

factory = StopWordRemoverFactory()
# 获取基础停用词集合
base_stopwords = factory.get_stop_words()
# 合并自定义停用词
merged_stopwords = set(base_stopwords + custom_stopwords)
# 创建包含自定义词的停用词移除器
stopword_remover = factory.create_stop_word_remover(merged_stopwords)

3. 调整文本处理顺序(关键)

如果停用词过滤步骤在小写转换/俚语替换之前,会导致匹配失败。正确的处理顺序应该是:

  1. 小写转换
  2. 俚语替换
  3. 停用词过滤
  4. 词干提取

修改后的clean_text示例:

def clean_text(text):
    # 1. 小写转换
    text = text.lower()
    # 2. 俚语替换(假设你有slang_dict映射字典)
    for slang, formal in slang_dict.items():
        text = text.replace(slang, formal)
    # 3. 停用词过滤
    text = stopword_remover.remove(text)
    # 4. 词干提取(假设你有已初始化的stemmer实例)
    words = text.split()
    stemmed_words = [stemmer.stem(word) for word in words]
    return ' '.join(stemmed_words)

4. 验证停用词集合

在代码里打印验证,确认目标停用词已被包含:

print("dok" in merged_stopwords)  # 正常应该返回True

测试效果

输入:Dok,anak saya sudah imunisasi DPT
修复后输出会移除"dok",得到类似anak saya sudah imunisasi dpt(词干提取后会进一步处理)的结果。

内容的提问来源于stack exchange,提问作者abbym

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 15:20:28