Sklearn添加自定义停用词无效:为何目标词仍出现在词频统计中?
问题:自定义停用词未生效,目标词汇仍出现在词频统计结果中
我正在测试基于以下代码的文本词频统计功能,尝试将'okay'、'yeah'、'thank'、'im'添加到停用词表中,但这些词仍出现在统计结果的图表里。
测试代码
import matplotlib.pyplot as plt from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS from collections import Counter import pandas as pd # 原代码遗漏pandas导入,补充完整 df_new = pd.DataFrame(['okay', 'yeah', 'thank', 'im']) stop_words = text.ENGLISH_STOP_WORDS.union(df_new) #stop_words w_counts = Counter(w for w in ' '.join(df['text_without_stopwords']).split() if w.lower() not in stop_words) df_words = pd.DataFrame.from_dict(w_counts, orient='index').reset_index() df_words.columns = ['word','count'] import seaborn as sns # 选取出现频率最高的25个词 d = df_words.nlargest(columns="count", n = 25) plt.figure(figsize=(20,5)) ax = sns.barplot(data=d, x= "word", y = "count") ax.set(ylabel = 'Count') plt.show()
问题原因及解决方案
1. 停用词集合引用错误
你导入的是from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS,但代码中错误使用了text.ENGLISH_STOP_WORDS——并没有导入名为text的模块,正确写法是直接调用ENGLISH_STOP_WORDS。
2. 自定义停用词格式不匹配
ENGLISH_STOP_WORDS.union()需要传入字符串集合或可迭代的字符串序列,但你把自定义停用词放在了pd.DataFrame中,DataFrame本身不是直接的字符串序列,导致这些词无法被正确加入停用词表。
修正后的关键代码
import pandas as pd from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS df_new = pd.DataFrame(['okay', 'yeah', 'thank', 'im']) # 将DataFrame中的自定义停用词转为字符串列表 custom_stopwords = df_new[0].tolist() # 统一转为小写,适配后续w.lower()的判断逻辑 custom_stopwords = [word.lower() for word in custom_stopwords] # 合并系统停用词和自定义停用词 stop_words = ENGLISH_STOP_WORDS.union(custom_stopwords) # 词频统计逻辑保持不变 w_counts = Counter(w for w in ' '.join(df['text_without_stopwords']).split() if w.lower() not in stop_words)
额外提示
如果文本中存在目标词汇的大小写变体(比如'Okay'、'YEAH'),将自定义停用词统一转为小写后,能和w.lower()的判断逻辑完全匹配,避免遗漏。
内容的提问来源于stack exchange,提问作者ASH
相关产品推荐
相关产品推荐

