如何让WordCloud将Excel多词字符串作为单个观测生成词云?
让WordCloud将每个Excel单元格内容作为整体生成词云的方法
你当前代码里用str(df)把整列转成字符串后传给generate(),WordCloud会默认按空格拆分所有文本,所以像"Mental health"这种多词单元格会被拆成两个单独的词。要解决这个问题,核心是直接统计每个单元格内容的出现频率,然后让WordCloud基于这个频率字典生成词云,而不是让它自己分词。
修改后的完整代码
import pandas as pd import numpy as np from PIL import Image from wordcloud import WordCloud, STOPWORDS import matplotlib.pyplot as plt # 读取数据并提取目标列 df = pd.read_csv(r"C:\Users\.......\jj.csv", encoding='utf8') outcome_series = df["Outcome"] # 统计每个单元格内容的出现频率,转成字典格式 freq_dict = outcome_series.value_counts().to_dict() # 加载掩码图片 our_mask = np.array(Image.open(r"C:\Users\.....\baby.png")) stopwords = set(STOPWORDS) # 初始化词云,改用generate_from_frequencies传入频率字典 wc = WordCloud( background_color="white", font_path='arial', colormap='Reds', random_state=1, repeat=True, collocations=False, max_words=150, stopwords=stopwords, mask=our_mask, contour_width=1, contour_color='Gray' ).generate_from_frequencies(freq_dict) # 绘制并展示词云 plt.imshow(wc, interpolation='bilinear') plt.axis('off') plt.show()
关键说明
outcome_series.value_counts()会自动统计每个单元格内容的出现次数,比如"Mental health"出现5次,就会生成键值对"Mental health":5generate_from_frequencies()方法会直接使用你提供的频率字典,不会对字典的键(即单元格内容)进行分词,确保多词内容作为整体出现在词云中- 注意给文件路径加上前缀
r,避免转义字符引发的路径错误
内容的提问来源于stack exchange,提问作者fary
相关产品推荐
相关产品推荐

