词云生成问题:低频二元组United State错误显示
Twitter词云异常问题修复方案
问题现象
基于爬取的Twitter数据生成词云时出现异常:
- 单词
states出现214次,state出现64次,二者合计频次远高于低频短语"United State"(仅1条推文包含) - 词云错误展示了"United State"二元组,未优先呈现高频单个词汇
问题根源
- 文本拼接错误:代码中使用
','.join(words)将处理后的单词用逗号连接,WordCloud的generate方法会对输入文本重新分词,逗号分隔的"united,state"会被误识别为短语"united state",导致该低频组合被统计。 - 未做词形归一化:
state和states属于同一语义的单复数形式,未合并词频,导致单个词的权重被分散,降低了在词云中的展示优先级。
修复代码
import numpy as np import matplotlib.pyplot as plt import re from PIL import Image from wordcloud import WordCloud from nltk.stem import WordNetLemmatizer import nltk nltk.download('wordnet') nltk.download('averaged_perceptron_tagger') STOPWORDS = [ 'i', 'me', 'my', 'myself', 'we', 'our', 'ours', 'ourselves', 'you', 'your', 'yours', 'yourself', 'yourselves', 'he', 'him', 'his', 'himself', 'she', 'her', 'hers', 'herself', 'it', 'its', 'itself', 'they', 'them', 'their', 'theirs', 'themselves', 'what', 'which', 'who', 'would', 'whom', 'this', 'that', 'these', 'those', 'am', 'is', 'are', 'was', 'were', 'be', 'been', 'being', 'have', 'has', 'had', 'having', 'do', 'does', 'did', 'doing', 'a', 'an', 'the', 'and', 'but', 'if', 'or', 'because', 'as', 'until', 'while', 'of', 'at', 'by', 'for', 'with', 'about', 'against', 'between', 'into', 'through', 'during', 'before', 'after', 'above', 'below', 'to', 'from', 'up', 'down', 'in', 'out', 'on', 'off', 'over', 'under', 'again', 'further', 'then', 'once', 'here', 'there', 'when', 'where', 'why', 'how', 'all', 'any', 'both', 'each', 'few', 'more', 'most', 'other', 'some', 'such', 'no', 'nor', 'not', 'only', 'own', 'same', 'so', 'than', 'too', 'very', 't', 'can', 'will', 'just', 'don', 'should', 'now' ] # 处理推文文本 raw_string = ' '.join(df['Tweet']) no_links = re.sub(r'http\S+', '', raw_string) no_unicode = re.sub(r"\\[a-z][a-z]?[0-9]+", '', no_links) no_special_characters = re.sub('[^A-Za-z ]+', '', no_unicode) words = no_special_characters.split(" ") words = [w.lower() for w in words if len(w) > 2 and w.lower() not in STOPWORDS] # 词形还原:将单复数、时态词统一为原形 lemmatizer = WordNetLemmatizer() lemmatized_words = [] for word in words: # 获取词性提升还原准确性 pos_tag = nltk.pos_tag([word])[0][1][0].lower() if pos_tag in ['n', 'v']: lemmatized_words.append(lemmatizer.lemmatize(word, pos=pos_tag)) else: lemmatized_words.append(lemmatizer.lemmatize(word)) # 统计词频 word_counts = {} for word in lemmatized_words: word_counts[word] = word_counts.get(word, 0) + 1 # 生成词云 mask = np.array(Image.open('Logo_location')) wc = WordCloud(background_color="white", max_words=2000, mask=mask, stopwords=STOPWORDS, relative_scaling=1) # 用词频字典生成,避免分词歧义 wc.generate_from_frequencies(word_counts) f = plt.figure(figsize=(13,13)) plt.imshow(wc, interpolation='bilinear') plt.title('Twitter Generated Cloud', size=30) plt.axis("off") plt.show()
关键修复点
- 替换
','.join(words)为空格拼接文本,或改用generate_from_frequencies传入词频字典,彻底避免分词歧义 - 添加词形还原步骤,将
states合并为state,集中词频权重 - 提前过滤停用词,减少无效词干扰
内容的提问来源于stack exchange,提问作者Rei Daemondheart
相关产品推荐
相关产品推荐

