You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

词云生成问题:低频二元组United State错误显示

Twitter词云异常问题修复方案

问题现象

基于爬取的Twitter数据生成词云时出现异常:

  • 单词states出现214次,state出现64次,二者合计频次远高于低频短语"United State"(仅1条推文包含)
  • 词云错误展示了"United State"二元组,未优先呈现高频单个词汇

问题根源

  1. 文本拼接错误:代码中使用','.join(words)将处理后的单词用逗号连接,WordCloud的generate方法会对输入文本重新分词,逗号分隔的"united,state"会被误识别为短语"united state",导致该低频组合被统计。
  2. 未做词形归一化:state和states属于同一语义的单复数形式,未合并词频,导致单个词的权重被分散,降低了在词云中的展示优先级。

修复代码

import numpy as np
import matplotlib.pyplot as plt
import re
from PIL import Image
from wordcloud import WordCloud
from nltk.stem import WordNetLemmatizer
import nltk
nltk.download('wordnet')
nltk.download('averaged_perceptron_tagger')

STOPWORDS = [
    'i', 'me', 'my', 'myself', 'we', 'our', 'ours', 'ourselves', 'you', 'your',
    'yours', 'yourself', 'yourselves', 'he', 'him', 'his', 'himself', 'she', 'her',
    'hers', 'herself', 'it', 'its', 'itself', 'they', 'them', 'their', 'theirs',
    'themselves', 'what', 'which', 'who', 'would', 'whom', 'this', 'that', 'these', 'those',
    'am', 'is', 'are', 'was', 'were', 'be', 'been', 'being', 'have', 'has', 'had',
    'having', 'do', 'does', 'did', 'doing', 'a', 'an', 'the', 'and', 'but', 'if',
    'or', 'because', 'as', 'until', 'while', 'of', 'at', 'by', 'for', 'with',
    'about', 'against', 'between', 'into', 'through', 'during', 'before', 'after',
    'above', 'below', 'to', 'from', 'up', 'down', 'in', 'out', 'on', 'off', 'over',
    'under', 'again', 'further', 'then', 'once', 'here', 'there', 'when', 'where',
    'why', 'how', 'all', 'any', 'both', 'each', 'few', 'more', 'most', 'other',
    'some', 'such', 'no', 'nor', 'not', 'only', 'own', 'same', 'so', 'than', 'too',
    'very', 't', 'can', 'will', 'just', 'don', 'should', 'now'
]

# 处理推文文本
raw_string = ' '.join(df['Tweet'])
no_links = re.sub(r'http\S+', '', raw_string)
no_unicode = re.sub(r"\\[a-z][a-z]?[0-9]+", '', no_links)
no_special_characters = re.sub('[^A-Za-z ]+', '', no_unicode)

words = no_special_characters.split(" ")
words = [w.lower() for w in words if len(w) > 2 and w.lower() not in STOPWORDS]

# 词形还原:将单复数、时态词统一为原形
lemmatizer = WordNetLemmatizer()
lemmatized_words = []
for word in words:
    # 获取词性提升还原准确性
    pos_tag = nltk.pos_tag([word])[0][1][0].lower()
    if pos_tag in ['n', 'v']:
        lemmatized_words.append(lemmatizer.lemmatize(word, pos=pos_tag))
    else:
        lemmatized_words.append(lemmatizer.lemmatize(word))

# 统计词频
word_counts = {}
for word in lemmatized_words:
    word_counts[word] = word_counts.get(word, 0) + 1

# 生成词云
mask = np.array(Image.open('Logo_location')) 
wc = WordCloud(background_color="white", max_words=2000, mask=mask, stopwords=STOPWORDS, relative_scaling=1)
# 用词频字典生成,避免分词歧义
wc.generate_from_frequencies(word_counts)

f = plt.figure(figsize=(13,13))
plt.imshow(wc, interpolation='bilinear')
plt.title('Twitter Generated Cloud', size=30)
plt.axis("off")
plt.show()

关键修复点

  • 替换','.join(words)为空格拼接文本,或改用generate_from_frequencies传入词频字典,彻底避免分词歧义
  • 添加词形还原步骤,将states合并为state,集中词频权重
  • 提前过滤停用词,减少无效词干扰

内容的提问来源于stack exchange,提问作者Rei Daemondheart

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 01:45:55