如何遍历记录并按分组生成多个WordCloud词云图?
按分组生成词云的解决方案
核心做法
要按Bin或Duration分组生成词云,关键是先按分组字段把对应文本(或带权重的统计结果)聚合好,再给每个分组单独生成词云。如果想用Duration来决定词的大小(比如时长越长的问题越显眼),得先统计每个分组里每个词的总时长,再用这个权重生成词云。
方法1:按Bin分组生成词云(看出现次数)
要是只想看每个Bin里词的出现次数,直接按Bin分组,把每个组的MyTExt拼起来,再生成词云就行:
import pandas as pd from wordcloud import WordCloud import matplotlib.pyplot as plt # 初始化数据 data = {'MyTExt':['MainOutage', 'MainOutage', 'Bills', 'Bills', 'Payments', 'Payments', 'Menu', 'Menu', 'Menu'], 'Duration':[200, 200, 400, 500, 20, 40, 50, 50, 60], 'Bin':['(23.6, 771.0]', '(23.6, 771.0]', '(23.6, 771.0]', '(23.6, 771.0]', '(771.0, 1511.0]', '(771.0, 1511.0]', '(771.0, 1511.0]', '(771.0, 1511.0]', '(771.0, 1511.0]']} df = pd.DataFrame(data) # 按Bin分组,把每个组的文本拼在一起 grouped_bin = df.groupby('Bin')['MyTExt'].apply(lambda x: ' '.join(x)) # 挨个分组生成词云 for bin_name, text in grouped_bin.items(): wordcloud = WordCloud(width=800, height=400, background_color='white').generate(text) plt.figure(figsize=(10,5)) plt.imshow(wordcloud) plt.title(f'Bin {bin_name} 的词云') plt.axis('off') plt.show()
方法2:按Bin分组生成词云(按Duration权重)
如果希望词的大小由这个词在分组里的总Duration决定(更贴合实际业务,比如时长久的问题更重要),可以先统计每个分组里每个词的总时长,再传给词云:
# 按Bin和MyTExt分组,算每个词的总Duration weighted_data = df.groupby(['Bin', 'MyTExt'])['Duration'].sum().unstack(fill_value=0) # 遍历每个Bin生成带权重的词云 for bin_name, weights in weighted_data.iterrows(): # 把统计结果转成字典,WordCloud支持直接用 weight_dict = weights.to_dict() # 过滤掉权重为0的词(可选,省得占资源) weight_dict = {k:v for k,v in weight_dict.items() if v>0} wordcloud = WordCloud(width=800, height=400, background_color='white').generate_from_frequencies(weight_dict) plt.figure(figsize=(10,5)) plt.imshow(wordcloud) plt.title(f'Bin {bin_name} 按时长加权的词云') plt.axis('off') plt.show()
方法3:按Duration分组生成词云
要是想直接按Duration分组,注意Duration是数值,直接按每个数值分的话会生成很多无意义的词云,所以建议先分箱(比如和现有Bin逻辑一致,或者自己定区间),再生成词云:
# 自定义Duration分箱(比如按100为间隔) df['Duration_Bin'] = pd.cut(df['Duration'], bins=[0, 100, 500, 2000], labels=['(0,100]', '(100,500]', '(500,2000]']) # 按自定义的Duration_Bin分组,用时长权重生成词云 weighted_duration = df.groupby(['Duration_Bin', 'MyTExt'])['Duration'].sum().unstack(fill_value=0) for bin_name, weights in weighted_duration.iterrows(): weight_dict = {k:v for k,v in weights.to_dict().items() if v>0} wordcloud = WordCloud(width=800, height=400, background_color='white').generate_from_frequencies(weight_dict) plt.figure(figsize=(10,5)) plt.imshow(wordcloud) plt.title(f'Duration区间 {bin_name} 的加权词云') plt.axis('off') plt.show()
注意点
- 只关注出现次数的话,直接拼接文本生成词云就够了;
- 要体现业务优先级的话,用
generate_from_frequencies()方法传权重字典更合适; - 按数值字段(比如Duration)分组时,一定要先分箱,不然每个单独数值都生成词云完全没用。
内容的提问来源于stack exchange,提问作者ASH
相关产品推荐
相关产品推荐

