You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Sklearn添加自定义停用词无效:为何目标词仍出现在词频统计中?

问题:自定义停用词未生效,目标词汇仍出现在词频统计结果中

我正在测试基于以下代码的文本词频统计功能,尝试将'okay'、'yeah'、'thank'、'im'添加到停用词表中,但这些词仍出现在统计结果的图表里。

测试代码

import matplotlib.pyplot as plt
from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS
from collections import Counter
import pandas as pd  # 原代码遗漏pandas导入,补充完整

df_new = pd.DataFrame(['okay', 'yeah', 'thank', 'im'])
stop_words = text.ENGLISH_STOP_WORDS.union(df_new)
#stop_words

w_counts = Counter(w for w in ' '.join(df['text_without_stopwords']).split() if w.lower() not in stop_words)


df_words = pd.DataFrame.from_dict(w_counts, orient='index').reset_index()
df_words.columns = ['word','count']


import seaborn as sns
# 选取出现频率最高的25个词
d = df_words.nlargest(columns="count", n = 25) 
plt.figure(figsize=(20,5))
ax = sns.barplot(data=d, x= "word", y = "count")
ax.set(ylabel = 'Count')
plt.show()

问题原因及解决方案

1. 停用词集合引用错误

你导入的是from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS,但代码中错误使用了text.ENGLISH_STOP_WORDS——并没有导入名为text的模块,正确写法是直接调用ENGLISH_STOP_WORDS。

2. 自定义停用词格式不匹配

ENGLISH_STOP_WORDS.union()需要传入字符串集合或可迭代的字符串序列,但你把自定义停用词放在了pd.DataFrame中,DataFrame本身不是直接的字符串序列,导致这些词无法被正确加入停用词表。

修正后的关键代码

import pandas as pd
from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS

df_new = pd.DataFrame(['okay', 'yeah', 'thank', 'im'])
# 将DataFrame中的自定义停用词转为字符串列表
custom_stopwords = df_new[0].tolist()
# 统一转为小写,适配后续w.lower()的判断逻辑
custom_stopwords = [word.lower() for word in custom_stopwords]
# 合并系统停用词和自定义停用词
stop_words = ENGLISH_STOP_WORDS.union(custom_stopwords)

# 词频统计逻辑保持不变
w_counts = Counter(w for w in ' '.join(df['text_without_stopwords']).split() if w.lower() not in stop_words)

额外提示

如果文本中存在目标词汇的大小写变体(比如'Okay'、'YEAH'),将自定义停用词统一转为小写后,能和w.lower()的判断逻辑完全匹配,避免遗漏。

内容的提问来源于stack exchange,提问作者ASH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 23:32:16