如何对文本预处理移除Gensim停用词 为何we、however等词未被过滤
问题根本原因
你代码的函数调用顺序错误是导致停用词漏删的核心原因:
- gensim内置的
remove_stopwords函数仅能匹配全小写的停用词,其内置停用词表所有条目均为小写格式 - 你当前的执行顺序是「先调用
remove_stopwords处理原始文本 → 再调用simple_preprocess做小写转换+分词」,原始文本里的We、However首字母为大写,调用停用词移除时无法匹配到小写的停用词条目,后续转小写已经错过了停用词过滤阶段,所以会残留。而the在原文本里本身就是小写,所以能被正常移除。
修复方案
方案1:调整函数调用顺序,分词后再过滤停用词(更推荐)
把分词和小写转换前置,再针对分词结果做停用词过滤,代码修改如下:
from gensim.parsing.preprocessing import STOPWORDS, simple_preprocess def read_text(text_path): text = [] with open(text_path) as file: lines = file.readlines() for line in lines: # 先做小写转换+分词,再过滤停用词 words = simple_preprocess(line) words_filtered = [w for w in words if w not in STOPWORDS] text.append(words_filtered) return text
方案2:先将整行文本转为小写,再调用remove_stopwords
如果要保留原有调用逻辑,提前做全局小写转换即可:
def read_text(text_path): text = [] with open(text_path) as file: lines = file.readlines() for index, line in enumerate(lines): # 先转小写再移除停用词 text.append(simple_preprocess(remove_stopwords(line.lower()))) return text
内容的提问来源于stack exchange,提问作者G. Macia
相关产品推荐
相关产品推荐

