You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于WordCloud词云未显示高频非停用词的技术咨询

解决词云未显示高频非停用词的问题

问题描述

使用WordCloud自带的STOPWORDS移除停用词后,词云未显示多个排名靠前的非停用词(如"data"、"learning"、"training"等,这些词均不在STOPWORDS列表中),但它们的出现频率高于已显示的"distribution"、"algorithm"等词。

相关代码:

from wordcloud import WordCloud, STOPWORDS
text = open("test.txt", mode="r", encoding="utf-8").read()

wc = WordCloud(background_color="white", stopwords=STOPWORDS, height=400, width=600)
wc.generate(text)
wc.to_file("my_first_word_cloud.png")

解决方案

问题核心是WordCloud默认的generate方法分词逻辑存在局限(仅按空格分词,未处理标点、大小写),导致高频词被拆分或识别错误,无法正确统计词频。以下是具体解决步骤:

1. 手动预处理文本与统计词频

先对文本进行清洗,确保分词准确,再手动统计词频,避免默认分词的问题:

from wordcloud import WordCloud, STOPWORDS
import re
from collections import Counter

# 读取原始文本
text = open("test.txt", mode="r", encoding="utf-8").read()

# 文本清洗:统一小写、移除标点符号
clean_text = re.sub(r'[^a-zA-Z\s]', '', text).lower()
# 按空格分词
words = clean_text.split()
# 过滤停用词
filtered_words = [word for word in words if word not in STOPWORDS]
# 统计词频
word_counts = Counter(filtered_words)

2. 基于词频字典生成词云

使用generate_from_frequencies方法传入手动统计的词频字典,确保高频词按实际频率显示:

# 初始化词云对象
wc = WordCloud(background_color="white", stopwords=STOPWORDS, height=400, width=600)
# 基于词频生成词云
wc.generate_from_frequencies(word_counts)
# 保存词云图片
wc.to_file("my_first_word_cloud.png")

3. 额外验证步骤

  • 确认目标词不在STOPWORDS中:执行print("data" in STOPWORDS),返回False则说明确实未被标记为停用词。
  • 查看统计后的词频:打印word_counts.most_common(10),确认"data"、"learning"等词的排名是否符合预期。

内容的提问来源于stack exchange,提问作者PythonistGal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 17:20:48