基于词频在matplotlib-set-diagrams中生成韦恩图词云的问题
韦恩图词云结合词频的优雅实现方案
核心思路
无需拆分三个独立蒙版,而是借助matplotlib-set-diagrams生成的韦恩图区域边界,结合wordcloud的词频权重逻辑,让词语按所属集合(仅词表1、仅词表2、交集)和自身词频,自动匹配字号并布局到对应区域。
具体实现步骤
- 提取韦恩图区域路径
- 用
matplotlib-set-diagrams绘制基础韦恩图,获取三个区域(仅A、仅B、交集)的路径对象,后续将这些路径转换为词云的布局边界。
- 用
- 拆分词频字典
- 把总词频数据拆分为三个子字典:
freq_only_A:仅出现在词表1的词语及其对应词频freq_only_B:仅出现在词表2的词语及其对应词频freq_both:同时存在于两个词表的词语及其词频(可根据需求取平均、总和或最大值)
- 把总词频数据拆分为三个子字典:
- 关联区域与词频布局
- 将每个区域的路径转换为词云掩码,调用
wordcloud的generate_from_frequencies方法,分别为三个词频字典生成词云,再将词云精准叠加到韦恩图的对应区域上。
- 将每个区域的路径转换为词云掩码,调用
简化代码示例
import matplotlib.pyplot as plt from matplotlib_set_diagrams import venn2 from wordcloud import WordCloud import numpy as np # 准备词频数据 freq_A = {"苹果": 10, "香蕉": 8, "橙子": 5} freq_B = {"香蕉": 12, "葡萄": 9, "橙子": 6} # 拆分词频集合 freq_only_A = {k: v for k, v in freq_A.items() if k not in freq_B} freq_only_B = {k: v for k, v in freq_B.items() if k not in freq_A} freq_both = {k: (freq_A[k] + freq_B[k])/2 for k in freq_A if k in freq_B} # 生成韦恩图并提取区域路径 fig, ax = plt.subplots() venn = venn2(subsets=(len(freq_only_A), len(freq_only_B), len(freq_both)), ax=ax) plt.close(fig) path_only_A = venn.get_patch_by_id('10').get_path() path_only_B = venn.get_patch_by_id('01').get_path() path_both = venn.get_patch_by_id('11').get_path() # 定义区域内词云生成函数 def create_wordcloud_in_path(freq, path, ax): x, y = path.vertices.T # 生成对应区域的掩码 mask = np.zeros((int(y.max()-y.min()), int(x.max()-x.min())), dtype=np.uint8) grid = np.mgrid[y.min():y.max(), x.min():x.max()].reshape(2, -1).T mask[grid[:,0].astype(int)-int(y.min()), grid[:,1].astype(int)-int(x.min())] = path.contains_points(grid) wc = WordCloud(background_color=None, mode="RGBA", mask=mask, relative_scaling=0.5) wc.generate_from_frequencies(freq) # 在指定区域绘制词云 ax.imshow(wc, extent=(x.min(), x.max(), y.min(), y.max()), aspect='auto') # 绘制最终韦恩图词云 fig, ax = plt.subplots(figsize=(8,8)) venn2(subsets=(len(freq_only_A), len(freq_only_B), len(freq_both)), ax=ax) create_wordcloud_in_path(freq_only_A, path_only_A, ax) create_wordcloud_in_path(freq_only_B, path_only_B, ax) create_wordcloud_in_path(freq_both, path_both, ax) plt.axis('off') plt.show()
优化方向
- 词频权重自定义:交集区域的词频可选择取两个集合的最大值、总和或平均值,适配不同可视化需求。
- 布局冲突解决:若区域内词语重叠,可调整
wordcloud的scale参数放大画布,或设置max_words限制显示数量。 - 视觉风格统一:为三个区域的词云设置相同字体、配色方案,保证整体视觉协调。
内容的提问来源于stack exchange,提问作者Simon Eaton
相关产品推荐
相关产品推荐

