You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对文本预处理移除Gensim停用词 为何we、however等词未被过滤

问题根本原因

你代码的函数调用顺序错误是导致停用词漏删的核心原因:

  • gensim内置的remove_stopwords函数仅能匹配全小写的停用词,其内置停用词表所有条目均为小写格式
  • 你当前的执行顺序是「先调用remove_stopwords处理原始文本 → 再调用simple_preprocess做小写转换+分词」,原始文本里的We、However首字母为大写,调用停用词移除时无法匹配到小写的停用词条目,后续转小写已经错过了停用词过滤阶段,所以会残留。而the在原文本里本身就是小写,所以能被正常移除。

修复方案

方案1:调整函数调用顺序,分词后再过滤停用词(更推荐)

把分词和小写转换前置,再针对分词结果做停用词过滤,代码修改如下:

from gensim.parsing.preprocessing import STOPWORDS, simple_preprocess

def read_text(text_path):
    text = []
    with open(text_path) as file:
        lines = file.readlines()
        for line in lines:
            # 先做小写转换+分词,再过滤停用词
            words = simple_preprocess(line)
            words_filtered = [w for w in words if w not in STOPWORDS]
            text.append(words_filtered)
    return text

方案2:先将整行文本转为小写,再调用remove_stopwords

如果要保留原有调用逻辑,提前做全局小写转换即可:

def read_text(text_path):
    text = []
    with open(text_path) as file:
        lines = file.readlines()
        for index, line in enumerate(lines):
            # 先转小写再移除停用词
            text.append(simple_preprocess(remove_stopwords(line.lower())))
    return text

内容的提问来源于stack exchange,提问作者G. Macia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 09:15:03