使用Matplotlib可视化主题模型时图表为空,请求代码排查
主题建模可视化图表为空的代码排查与修复
问题描述
开展主题建模工作,使用Matplotlib将词频与关键词权重在同一张图表中展示,DataFrame已包含对应数据,但生成的图表为空,需排查代码问题。
原始代码
from collections import Counter topics = final_model.show_topics(formatted=False) data_flat = [w for w_list in Tweet_df.tokens for w in w_list] counter = Counter(data_flat) out = [] for i, topic in topics: for word, weight in topic: out.append([word, i , weight, counter[word]]) df = pd.DataFrame(out, columns=['word', 'topic_id', 'importance', 'word_count']) print(df) # Plot Word Count and Weights of Topic Keywords fig, axes = plt.subplots(2, 2, figsize=(16,10), sharey=True, dpi=160) cols = [color for name, color in mcolors.TABLEAU_COLORS.items()] for i, ax in enumerate(axes.flatten()): ax.bar(x='word', height="word_count", data=df.loc[df.topic_id==i, :], color=cols[i], width=0.5, alpha=0.3, label='Word Count') ax_twin = ax.twinx() ax_twin.bar(x='word', height="importance", data=df.loc[df.topic_id==i, :], color=cols[i], width=0.2, label='Weights') ax.set_ylabel('Word Count', color=cols[i]) ax_twin.set_ylim(0, 0.030); ax.set_ylim(0, 3500) ax.set_title('Topic: ' + str(i), color=cols[i], fontsize=16) ax.tick_params(axis='y', left=False) ax.set_xticklabels(df.loc[df.topic_id==i, 'word'], rotation=30, horizontalalignment= 'right') ax.legend(loc='upper left'); ax_twin.legend(loc='upper right') fig.tight_layout(w_pad=2) fig.suptitle('Word Count and Importance of Topic Keywords', fontsize=22, y=1.05) plt.show()
核心问题与修复方案
子图数量与实际主题数不匹配:代码固定创建2×2=4个子图,但如果模型生成的主题数不是4个,循环到超出实际
topic_id的索引时,df.loc[df.topic_id==i, :]会返回空DataFrame,对应子图无内容。
修复:先获取主题数量,动态计算子图行列数,只遍历存在的主题:import math num_topics = len(topics) cols_num = 2 rows_num = math.ceil(num_topics / cols_num) fig, axes = plt.subplots(rows_num, cols_num, figsize=(16,10), sharey=True, dpi=160) for i, ax in enumerate(axes.flatten()[:num_topics]): topic_data = df.loc[df.topic_id==i, :] # 后续绘图逻辑set_xticklabels使用错误:直接设置标签但未指定刻度位置,导致标签无法绑定到bar图位置,引发图表无内容。
修复:先设置x轴刻度位置,再绑定标签:topic_words = topic_data['word'].tolist() ax.set_xticks(range(len(topic_words))) ax.set_xticklabels(topic_words, rotation=30, horizontalalignment='right')双轴bar图x轴重叠+数据绑定问题:使用
x='word'作为分类轴,双轴bar完全重叠,且可能因分类轴处理问题导致图表不显示。
修复:改用数值型x轴,给第二个bar偏移位置避免重叠:import numpy as np x = np.arange(len(topic_words)) ax.bar(x=x, height=topic_data['word_count'], width=0.35, alpha=0.3, label='Word Count') ax_twin.bar(x=x+0.35, height=topic_data['importance'], width=0.35, label='Weights')轴范围硬编码适配性差:硬设置的
ylim如果和实际数据范围不匹配,会导致bar图被截断或显得极小。
修复:动态设置y轴范围,或让Matplotlib自动适配:ax.set_ylim(bottom=0, top=topic_data['word_count'].max()*1.1) ax_twin.set_ylim(bottom=0, top=topic_data['importance'].max()*1.1)
完整修复后的代码示例
from collections import Counter import math import numpy as np import matplotlib.pyplot as plt import matplotlib.colors as mcolors import pandas as pd topics = final_model.show_topics(formatted=False) data_flat = [w for w_list in Tweet_df.tokens for w in w_list] counter = Counter(data_flat) out = [] for i, topic in topics: for word, weight in topic: out.append([word, i , weight, counter[word]]) df = pd.DataFrame(out, columns=['word', 'topic_id', 'importance', 'word_count']) print(df) # Plot Word Count and Weights of Topic Keywords num_topics = len(topics) cols_num = 2 rows_num = math.ceil(num_topics / cols_num) fig, axes = plt.subplots(rows_num, cols_num, figsize=(16,10), sharey=True, dpi=160) cols = [color for name, color in mcolors.TABLEAU_COLORS.items()] # 遍历每个存在的主题 for i, ax in enumerate(axes.flatten()[:num_topics]): topic_data = df.loc[df.topic_id==i, :] topic_words = topic_data['word'].tolist() x = np.arange(len(topic_words)) # 绘制词频bar ax.bar(x=x, height=topic_data['word_count'], color=cols[i], width=0.35, alpha=0.3, label='Word Count') ax_twin = ax.twinx() # 绘制权重bar,偏移x轴位置避免重叠 ax_twin.bar(x=x+0.35, height=topic_data['importance'], color=cols[i], width=0.35, label='Weights') ax.set_ylabel('Word Count', color=cols[i]) # 动态设置y轴范围 ax.set_ylim(bottom=0, top=topic_data['word_count'].max()*1.1) ax_twin.set_ylim(bottom=0, top=topic_data['importance'].max()*1.1) ax.set_title('Topic: ' + str(i), color=cols[i], fontsize=16) ax.tick_params(axis='y', left=False) # 让标签居中在两个bar中间 ax.set_xticks(x + 0.175) ax.set_xticklabels(topic_words, rotation=30, horizontalalignment='right') ax.legend(loc='upper left'); ax_twin.legend(loc='upper right') # 隐藏多余的子图(如果主题数不是cols_num的整数倍) for ax in axes.flatten()[num_topics:]: ax.axis('off') fig.tight_layout(w_pad=2) fig.suptitle('Word Count and Importance of Topic Keywords', fontsize=22, y=1.05) plt.show()
内容的提问来源于stack exchange,提问作者socialscientist90
相关产品推荐
相关产品推荐

