You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中修改NLTK停用词列表,保留this/these?

调整NLTK停用词集合以保留指定词汇的简洁方案

嘿,完全不用不好意思——新手阶段的问题都是成长的必经之路😉!针对你的需求,其实不用手动罗列所有停用词,利用Python集合的特性就能优雅地解决:

步骤1:准备NLTK停用词

首先确保你已经下载了NLTK的停用词数据集(第一次使用需要执行):

import nltk
nltk.download('stopwords')

步骤2:创建自定义停用词集合

NLTK默认的停用词是一个集合,我们可以直接用集合差集操作移除你想保留的词汇(也就是把"this"和"these"从停用词列表中剔除,这样过滤时就不会删掉它们):

from nltk.corpus import stopwords

# 获取英文停用词集合
default_stop_words = set(stopwords.words('english'))

# 定义需要保留的词汇
words_to_retain = {"this", "these"}

# 生成自定义停用词集合:默认集合 - 要保留的词汇
custom_stop_words = default_stop_words - words_to_retain

验证效果(可选)

你可以用一段示例文本测试过滤效果,确认"this"和"these"被保留:

sample_sentence = "This is a test sentence, these words should stay while others get filtered."
# 分词(这里简化处理,实际场景建议用更专业的分词工具)
tokens = sample_sentence.lower().split()
# 过滤停用词
filtered_tokens = [token for token in tokens if token not in custom_stop_words]

print(filtered_tokens)
# 输出:['this', 'test', 'sentence,', 'these', 'words', 'stay', 'filtered.']

为什么这是优雅的方案?

  • 集合的差集操作-非常简洁,一行代码就能完成修改,不用手动遍历或罗列所有停用词
  • 集合的成员查询效率是O(1),比列表快得多,处理大体积CSV文件时性能更优

内容的提问来源于stack exchange,提问作者Paul Kremershof

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:38:29