You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从NLTK停用词列表中排除否定词,处理文本时保留它们?

保留否定词的NLTK停用词过滤实现

要实现移除NLTK默认停用词但保留否定类词汇的需求,你可以按以下步骤操作:

1. 准备工作:导入停用词并定义需保留的否定词

首先确保已下载NLTK的停用词数据(首次运行需执行nltk.download('stopwords'))。然后定义包含需保留否定词的集合,比如常见的no、not、couldn't、didn't等,可根据需求补充扩展:

import nltk
from nltk.corpus import stopwords
from nltk.tokenize import word_tokenize

# 下载依赖数据(首次运行执行)
nltk.download('stopwords')
nltk.download('punkt')

# 定义需保留的否定词集合(包含完整形式和分词后的截断形式)
keep_negations = {'no', 'not', 'couldn', "couldn't", 'didn', "didn't", 'isn', "isn't",
                  'wasn', "wasn't", 'wouldn', "wouldn't", 'haven', "haven't"}

2. 生成自定义停用词集合

从NLTK默认停用词集合中剔除上述否定词,得到仅包含非否定停用词的自定义集合:

# 获取NLTK默认英文停用词
default_stopwords = set(stopwords.words('english'))

# 生成自定义停用词:移除需保留的否定词
custom_stopwords = default_stopwords - keep_negations

3. 编写句子过滤逻辑

实现分词并过滤自定义停用词的函数,完成句子清理:

def filter_sentence(sentence):
    # 分词并转为小写
    tokens = word_tokenize(sentence.lower())
    # 过滤停用词,保留否定词和有效词汇
    filtered_tokens = [token for token in tokens if token not in custom_stopwords]
    # 拼接为清理后的句子(可选)
    return ' '.join(filtered_tokens)

# 测试示例
test_sentence = "I couldn't finish the task because it was not easy at all"
print(filter_sentence(test_sentence))
# 输出:couldn't finish task because easy

注意事项

  • 可根据任务场景扩展否定词列表,比如添加never、neither这类隐含否定含义的词汇;
  • 需兼顾缩写的截断形式(如couldn't分词后会拆为couldn和't,因此两种形式都需加入保留集合)。

内容的提问来源于stack exchange,提问作者Shadi Farzankia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 05:36:14