如何生成带黑边的圆形LDA主题词云并应用于PCA图表
LDA主题词云定制实现方案
需求说明
适配出版规范,为LDA生成的每个主题制作圆形、带黑色边框的词云图,最终可将生成的词云应用到PCA图表中。
核心修改点
- 新增圆形掩膜参数限制词云生成范围,实现圆形效果
- 子图渲染时添加与词云同尺寸的圆形黑色边框,符合出版规范
- 保留词云对象输出逻辑,支持后续嵌入PCA图表
完整可运行代码
import pandas as pd import seaborn as sns import matplotlib.pyplot as plt import numpy as np from wordcloud import WordCloud sns.set() # 输入数据 topics_df = pd.DataFrame({ 'Unnamed: 0': {0: 'Topic1', 1: 'Topic2', 2: 'Topic3', 3: 'Topic4', 4: 'Topic5'}, 'Terms per Topic': { 0: 'charge, model, infrastructure, station, time, base, propose, drive, public, service, range, problem, study, paper, network, user, travel, battery, city, method', 1: 'cost, emission, fuel, car, high, price, low, reduce, hybrid, carbon, battery, gas, alternative, benefit, compare, find, conventional, efficiency, gasoline, total', 2: 'market, technology, policy, paper, change, development, support, mobility, business, industry, government, focus, study, transition, innovation, decision, develop, technological, case, lead', 3: 'consumer, policy, adoption, effect, model, purchase, bev, factor, evs, study, incentive, choice, preference, subsidy, influence, environmental, state, level, suggest, analysis', 4: 'energy, system, electricity, demand, transport, scenario, impact, increase, power, phev, transportation, sector, potential, base, show, consumption, reserve, term, economic, grid' } }) # 生成圆形掩膜:400*400像素正方形,中间圆形区域为可生成词云的区域 x, y = np.ogrid[:400, :400] mask = (x - 200) ** 2 + (y - 200) ** 2 > 190 ** 2 mask = 255 * mask.astype(int) # 词云初始化新增mask参数 wc = WordCloud( background_color="white", colormap="Dark2", max_font_size=150, random_state=42, mask=mask, contour_width=0 # 禁用自带边框,自定义绘制更灵活 ) plt.rcParams['figure.figsize'] = [20, 15] # 存储所有主题词云对象,后续嵌入PCA时可直接调用 topic_wc_list = [] # 创建每个主题的子图 for i in range(5): wc.generate(text=topics_df["Terms per Topic"][i]) topic_wc_list.append(wc.to_array()) # 保存词云数组 plt.subplot(5, 4, i+1) plt.imshow(wc, interpolation="bilinear") # 绘制黑色圆形边框,和词云大小匹配 ax = plt.gca() circle = plt.Circle((200, 200), 190, color='black', fill=False, linewidth=2) ax.add_patch(circle) plt.axis("off") plt.title(topics_df["Unnamed: 0"][i], fontsize=12) plt.show()
PCA嵌入参考逻辑
如果需要将生成的词云放到PCA图表对应位置,可参考以下步骤实现:
- 提前计算所有主题的PCA降维后坐标,存储为二维数组
topic_pca_coords - 导入相关工具类:
from matplotlib.offsetbox import OffsetImage, AnnotationBbox
- 在PCA图表生成后,遍历坐标和词云数组,绑定到对应位置:
# pca_ax为PCA图表的坐标轴对象 for idx, (x_coord, y_coord) in enumerate(topic_pca_coords): img = OffsetImage(topic_wc_list[idx], zoom=0.2) # 可调整zoom参数控制词云大小 ab = AnnotationBbox(img, (x_coord, y_coord), frameon=False) pca_ax.add_artist(ab)
内容的提问来源于stack exchange,提问作者Bloxx
相关产品推荐
相关产品推荐

