如何从NLTK停用词列表中排除否定词,处理文本时保留它们?
保留否定词的NLTK停用词过滤实现
要实现移除NLTK默认停用词但保留否定类词汇的需求,你可以按以下步骤操作:
1. 准备工作:导入停用词并定义需保留的否定词
首先确保已下载NLTK的停用词数据(首次运行需执行nltk.download('stopwords'))。然后定义包含需保留否定词的集合,比如常见的no、not、couldn't、didn't等,可根据需求补充扩展:
import nltk from nltk.corpus import stopwords from nltk.tokenize import word_tokenize # 下载依赖数据(首次运行执行) nltk.download('stopwords') nltk.download('punkt') # 定义需保留的否定词集合(包含完整形式和分词后的截断形式) keep_negations = {'no', 'not', 'couldn', "couldn't", 'didn', "didn't", 'isn', "isn't", 'wasn', "wasn't", 'wouldn', "wouldn't", 'haven', "haven't"}
2. 生成自定义停用词集合
从NLTK默认停用词集合中剔除上述否定词,得到仅包含非否定停用词的自定义集合:
# 获取NLTK默认英文停用词 default_stopwords = set(stopwords.words('english')) # 生成自定义停用词:移除需保留的否定词 custom_stopwords = default_stopwords - keep_negations
3. 编写句子过滤逻辑
实现分词并过滤自定义停用词的函数,完成句子清理:
def filter_sentence(sentence): # 分词并转为小写 tokens = word_tokenize(sentence.lower()) # 过滤停用词,保留否定词和有效词汇 filtered_tokens = [token for token in tokens if token not in custom_stopwords] # 拼接为清理后的句子(可选) return ' '.join(filtered_tokens) # 测试示例 test_sentence = "I couldn't finish the task because it was not easy at all" print(filter_sentence(test_sentence)) # 输出:couldn't finish task because easy
注意事项
- 可根据任务场景扩展否定词列表,比如添加
never、neither这类隐含否定含义的词汇; - 需兼顾缩写的截断形式(如
couldn't分词后会拆为couldn和't,因此两种形式都需加入保留集合)。
内容的提问来源于stack exchange,提问作者Shadi Farzankia
相关产品推荐
相关产品推荐

