遍历字典列表值移除停用词无效果,求问题原因与解决方法
问题原因与解决方案
错误根源
你的代码逻辑是将完整的短语字符串(比如'change my account')直接和停用词集合里的单个单词(比如'my')做对比,两者完全不是同一维度的内容,自然永远匹配不上,所以原字典不会有任何变化。
修正代码
要实现移除短语中的停用词,需要先把每个短语拆分为单个单词,过滤掉停用词后再重新拼接成短语:
from nltk.corpus import stopwords stop_words = set(stopwords.words("english")) data = {'ACCOUNT_CLOSURE': ['account closure', 'close account', 'close bank', 'terminate account', 'account deletion', 'cancel account', 'account cancellation'], 'ACCOUNT_CHANGE': ['change my account', 'switch my account', 'change from private into savings', 'convert into family package', 'change title of the account', 'make title account to family', 'help me access the documentation']} for key, phrases in data.items(): cleaned_phrases = [] for phrase in phrases: # 拆分短语为单词,过滤停用词,再重新拼接 cleaned_words = [word for word in phrase.split() if word.lower() not in stop_words] # 过滤后如果还有单词,就拼接成短语;否则跳过空短语 if cleaned_words: cleaned_phrases.append(' '.join(cleaned_words)) data[key] = cleaned_phrases print(data)
补充说明
- 加上
word.lower()是为了兼容短语中首字母大写的单词,避免因为大小写不匹配漏过滤停用词 - 如果某个短语过滤后所有单词都是停用词,可根据需求选择保留空字符串或直接跳过,示例代码选择跳过这类空短语
内容的提问来源于stack exchange,提问作者Wiliam
相关产品推荐
相关产品推荐

