如何在Python中修改NLTK停用词列表,保留this/these?
调整NLTK停用词集合以保留指定词汇的简洁方案
嘿,完全不用不好意思——新手阶段的问题都是成长的必经之路😉!针对你的需求,其实不用手动罗列所有停用词,利用Python集合的特性就能优雅地解决:
步骤1:准备NLTK停用词
首先确保你已经下载了NLTK的停用词数据集(第一次使用需要执行):
import nltk nltk.download('stopwords')
步骤2:创建自定义停用词集合
NLTK默认的停用词是一个集合,我们可以直接用集合差集操作移除你想保留的词汇(也就是把"this"和"these"从停用词列表中剔除,这样过滤时就不会删掉它们):
from nltk.corpus import stopwords # 获取英文停用词集合 default_stop_words = set(stopwords.words('english')) # 定义需要保留的词汇 words_to_retain = {"this", "these"} # 生成自定义停用词集合:默认集合 - 要保留的词汇 custom_stop_words = default_stop_words - words_to_retain
验证效果(可选)
你可以用一段示例文本测试过滤效果,确认"this"和"these"被保留:
sample_sentence = "This is a test sentence, these words should stay while others get filtered." # 分词(这里简化处理,实际场景建议用更专业的分词工具) tokens = sample_sentence.lower().split() # 过滤停用词 filtered_tokens = [token for token in tokens if token not in custom_stop_words] print(filtered_tokens) # 输出:['this', 'test', 'sentence,', 'these', 'words', 'stay', 'filtered.']
为什么这是优雅的方案?
- 集合的差集操作
-非常简洁,一行代码就能完成修改,不用手动遍历或罗列所有停用词 - 集合的成员查询效率是O(1),比列表快得多,处理大体积CSV文件时性能更优
内容的提问来源于stack exchange,提问作者Paul Kremershof
相关产品推荐
相关产品推荐

