You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Jupyter Notebook中从词云移除自定义词汇

移除词云中自定义无关词汇的解决方案

方法1:预处理文本时过滤自定义词汇

先定义需要移除的无关词汇列表,在生成词云前对文本做过滤:

# 自定义需要移除的词汇
custom_exclude = {'user', 'need', 'anyone'}

# 批量过滤文本内容
processed_texts = []
for text in df1['text'].tolist():
    # 统一转小写避免大小写遗漏,拆分后过滤词汇
    cleaned_words = [word for word in text.lower().split() if word not in custom_exclude]
    processed_texts.append(' '.join(cleaned_words))

# 生成词云
wordcloud = WordCloud(background_color="white", width=1600, height=800).generate(' '.join(processed_texts))
plt.figure(figsize=(20,10), facecolor='k')
plt.imshow(wordcloud)
plt.axis('off')  # 隐藏坐标轴优化展示效果
plt.show()

方法2:直接给WordCloud传入自定义停用词

WordCloud自带stopwords参数,可结合默认停用词扩充自定义列表:

from wordcloud import STOPWORDS

# 合并默认停用词和自定义排除词汇
combined_stopwords = STOPWORDS.union({'user', 'need', 'anyone'})

# 生成词云时指定停用词
wordcloud = WordCloud(
    background_color="white",
    width=1600,
    height=800,
    stopwords=combined_stopwords
).generate(' '.join(df1['text'].tolist()))

plt.figure(figsize=(20,10), facecolor='k')
plt.imshow(wordcloud)
plt.axis('off')
plt.show()

补充建议

  • 若文本存在大小写不一致的情况,务必统一转小写,防止漏过滤目标词汇
  • 可根据生成的词云结果,持续扩充自定义排除列表,逐步让词云聚焦到email、server、outlook等核心主题上

内容的提问来源于stack exchange,提问作者ChelseaSupencheck

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 00:45:24