You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Matplotlib可视化主题模型时图表为空,请求代码排查

主题建模可视化图表为空的代码排查与修复

问题描述

开展主题建模工作,使用Matplotlib将词频与关键词权重在同一张图表中展示,DataFrame已包含对应数据,但生成的图表为空,需排查代码问题。

原始代码

from collections import Counter
topics = final_model.show_topics(formatted=False)
data_flat = [w for w_list in Tweet_df.tokens for w in w_list]
counter = Counter(data_flat)

out = []
for i, topic in topics:
   for word, weight in topic:
       out.append([word, i , weight, counter[word]])

df = pd.DataFrame(out, columns=['word', 'topic_id', 'importance', 'word_count'])     

print(df)   

# Plot Word Count and Weights of Topic Keywords
fig, axes = plt.subplots(2, 2, figsize=(16,10), sharey=True, dpi=160)
cols = [color for name, color in mcolors.TABLEAU_COLORS.items()]
for i, ax in enumerate(axes.flatten()):
   ax.bar(x='word', height="word_count", data=df.loc[df.topic_id==i, :], color=cols[i], width=0.5, alpha=0.3, label='Word Count')
   ax_twin = ax.twinx()
   ax_twin.bar(x='word', height="importance", data=df.loc[df.topic_id==i, :], color=cols[i], width=0.2, label='Weights')
   ax.set_ylabel('Word Count', color=cols[i])
   ax_twin.set_ylim(0, 0.030); ax.set_ylim(0, 3500)
   ax.set_title('Topic: ' + str(i), color=cols[i], fontsize=16)
   ax.tick_params(axis='y', left=False)
   ax.set_xticklabels(df.loc[df.topic_id==i, 'word'], rotation=30, horizontalalignment= 'right')
   ax.legend(loc='upper left'); ax_twin.legend(loc='upper right')

fig.tight_layout(w_pad=2)
fig.suptitle('Word Count and Importance of Topic Keywords', fontsize=22, y=1.05)    
plt.show()

核心问题与修复方案

  • 子图数量与实际主题数不匹配:代码固定创建2×2=4个子图,但如果模型生成的主题数不是4个,循环到超出实际topic_id的索引时,df.loc[df.topic_id==i, :]会返回空DataFrame,对应子图无内容。
    修复:先获取主题数量,动态计算子图行列数,只遍历存在的主题:

    import math
    num_topics = len(topics)
    cols_num = 2
    rows_num = math.ceil(num_topics / cols_num)
    fig, axes = plt.subplots(rows_num, cols_num, figsize=(16,10), sharey=True, dpi=160)
    for i, ax in enumerate(axes.flatten()[:num_topics]):
        topic_data = df.loc[df.topic_id==i, :]
        # 后续绘图逻辑
    
  • set_xticklabels使用错误:直接设置标签但未指定刻度位置,导致标签无法绑定到bar图位置,引发图表无内容。
    修复:先设置x轴刻度位置,再绑定标签:

    topic_words = topic_data['word'].tolist()
    ax.set_xticks(range(len(topic_words)))
    ax.set_xticklabels(topic_words, rotation=30, horizontalalignment='right')
    
  • 双轴bar图x轴重叠+数据绑定问题:使用x='word'作为分类轴,双轴bar完全重叠,且可能因分类轴处理问题导致图表不显示。
    修复:改用数值型x轴,给第二个bar偏移位置避免重叠:

    import numpy as np
    x = np.arange(len(topic_words))
    ax.bar(x=x, height=topic_data['word_count'], width=0.35, alpha=0.3, label='Word Count')
    ax_twin.bar(x=x+0.35, height=topic_data['importance'], width=0.35, label='Weights')
    
  • 轴范围硬编码适配性差:硬设置的ylim如果和实际数据范围不匹配,会导致bar图被截断或显得极小。
    修复:动态设置y轴范围,或让Matplotlib自动适配:

    ax.set_ylim(bottom=0, top=topic_data['word_count'].max()*1.1)
    ax_twin.set_ylim(bottom=0, top=topic_data['importance'].max()*1.1)
    

完整修复后的代码示例

from collections import Counter
import math
import numpy as np
import matplotlib.pyplot as plt
import matplotlib.colors as mcolors
import pandas as pd

topics = final_model.show_topics(formatted=False)
data_flat = [w for w_list in Tweet_df.tokens for w in w_list]
counter = Counter(data_flat)

out = []
for i, topic in topics:
   for word, weight in topic:
       out.append([word, i , weight, counter[word]])

df = pd.DataFrame(out, columns=['word', 'topic_id', 'importance', 'word_count'])     

print(df)   

# Plot Word Count and Weights of Topic Keywords
num_topics = len(topics)
cols_num = 2
rows_num = math.ceil(num_topics / cols_num)
fig, axes = plt.subplots(rows_num, cols_num, figsize=(16,10), sharey=True, dpi=160)
cols = [color for name, color in mcolors.TABLEAU_COLORS.items()]

# 遍历每个存在的主题
for i, ax in enumerate(axes.flatten()[:num_topics]):
    topic_data = df.loc[df.topic_id==i, :]
    topic_words = topic_data['word'].tolist()
    x = np.arange(len(topic_words))
    
    # 绘制词频bar
    ax.bar(x=x, height=topic_data['word_count'], color=cols[i], width=0.35, alpha=0.3, label='Word Count')
    ax_twin = ax.twinx()
    # 绘制权重bar,偏移x轴位置避免重叠
    ax_twin.bar(x=x+0.35, height=topic_data['importance'], color=cols[i], width=0.35, label='Weights')
    
    ax.set_ylabel('Word Count', color=cols[i])
    # 动态设置y轴范围
    ax.set_ylim(bottom=0, top=topic_data['word_count'].max()*1.1)
    ax_twin.set_ylim(bottom=0, top=topic_data['importance'].max()*1.1)
    
    ax.set_title('Topic: ' + str(i), color=cols[i], fontsize=16)
    ax.tick_params(axis='y', left=False)
    # 让标签居中在两个bar中间
    ax.set_xticks(x + 0.175)
    ax.set_xticklabels(topic_words, rotation=30, horizontalalignment='right')
    
    ax.legend(loc='upper left'); ax_twin.legend(loc='upper right')

# 隐藏多余的子图(如果主题数不是cols_num的整数倍)
for ax in axes.flatten()[num_topics:]:
    ax.axis('off')

fig.tight_layout(w_pad=2)
fig.suptitle('Word Count and Importance of Topic Keywords', fontsize=22, y=1.05)    
plt.show()

内容的提问来源于stack exchange,提问作者socialscientist90

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 04:50:23