You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于词频在matplotlib-set-diagrams中生成韦恩图词云的问题

韦恩图词云结合词频的优雅实现方案

核心思路

无需拆分三个独立蒙版,而是借助matplotlib-set-diagrams生成的韦恩图区域边界,结合wordcloud的词频权重逻辑,让词语按所属集合(仅词表1、仅词表2、交集)和自身词频,自动匹配字号并布局到对应区域。

具体实现步骤

  1. 提取韦恩图区域路径
    • 用matplotlib-set-diagrams绘制基础韦恩图,获取三个区域(仅A、仅B、交集)的路径对象,后续将这些路径转换为词云的布局边界。
  2. 拆分词频字典
    • 把总词频数据拆分为三个子字典:
      • freq_only_A:仅出现在词表1的词语及其对应词频
      • freq_only_B:仅出现在词表2的词语及其对应词频
      • freq_both:同时存在于两个词表的词语及其词频(可根据需求取平均、总和或最大值)
  3. 关联区域与词频布局
    • 将每个区域的路径转换为词云掩码,调用wordcloud的generate_from_frequencies方法,分别为三个词频字典生成词云,再将词云精准叠加到韦恩图的对应区域上。

简化代码示例

import matplotlib.pyplot as plt
from matplotlib_set_diagrams import venn2
from wordcloud import WordCloud
import numpy as np

# 准备词频数据
freq_A = {"苹果": 10, "香蕉": 8, "橙子": 5}
freq_B = {"香蕉": 12, "葡萄": 9, "橙子": 6}

# 拆分词频集合
freq_only_A = {k: v for k, v in freq_A.items() if k not in freq_B}
freq_only_B = {k: v for k, v in freq_B.items() if k not in freq_A}
freq_both = {k: (freq_A[k] + freq_B[k])/2 for k in freq_A if k in freq_B}

# 生成韦恩图并提取区域路径
fig, ax = plt.subplots()
venn = venn2(subsets=(len(freq_only_A), len(freq_only_B), len(freq_both)), ax=ax)
plt.close(fig)

path_only_A = venn.get_patch_by_id('10').get_path()
path_only_B = venn.get_patch_by_id('01').get_path()
path_both = venn.get_patch_by_id('11').get_path()

# 定义区域内词云生成函数
def create_wordcloud_in_path(freq, path, ax):
    x, y = path.vertices.T
    # 生成对应区域的掩码
    mask = np.zeros((int(y.max()-y.min()), int(x.max()-x.min())), dtype=np.uint8)
    grid = np.mgrid[y.min():y.max(), x.min():x.max()].reshape(2, -1).T
    mask[grid[:,0].astype(int)-int(y.min()), grid[:,1].astype(int)-int(x.min())] = path.contains_points(grid)
    
    wc = WordCloud(background_color=None, mode="RGBA", mask=mask, relative_scaling=0.5)
    wc.generate_from_frequencies(freq)
    # 在指定区域绘制词云
    ax.imshow(wc, extent=(x.min(), x.max(), y.min(), y.max()), aspect='auto')

# 绘制最终韦恩图词云
fig, ax = plt.subplots(figsize=(8,8))
venn2(subsets=(len(freq_only_A), len(freq_only_B), len(freq_both)), ax=ax)
create_wordcloud_in_path(freq_only_A, path_only_A, ax)
create_wordcloud_in_path(freq_only_B, path_only_B, ax)
create_wordcloud_in_path(freq_both, path_both, ax)
plt.axis('off')
plt.show()

优化方向

  • 词频权重自定义:交集区域的词频可选择取两个集合的最大值、总和或平均值,适配不同可视化需求。
  • 布局冲突解决:若区域内词语重叠,可调整wordcloud的scale参数放大画布,或设置max_words限制显示数量。
  • 视觉风格统一:为三个区域的词云设置相同字体、配色方案,保证整体视觉协调。

内容的提问来源于stack exchange,提问作者Simon Eaton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 22:03:21