You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何结合Gensim LDA与TSNE绘制带词频权重的词标签聚类图?

改造LDA主题可视化:词标签替代散点并按词频调整大小

以下是修改后的完整代码,实现用词标签替代原散点图,并且词的大小与该词在语料库中的出现频次成正比:

from gensim.models import LdaModel
from gensim.corpora.dictionary import Dictionary
from sklearn.manifold import TSNE
from bokeh.plotting import figure, output_notebook, show
from bokeh.models import ColumnDataSource, LabelSet
import pandas as pd
import numpy as np
import matplotlib.colors as mcolors

# --- 你的原有LDA模型代码(保留即可)---
# dictionary = Dictionary(all_texts)
# corpus = [dictionary.doc2bow(text) for text in all_texts]
# lda_model = LdaModel(corpus=corpus, id2word=dictionary, num_topics=10, alpha="auto",per_word_topics=True,random_state=42)

# 1. 获取每个词的主题分布和词频
word_topic_dist = []
word_freq = []
words = []

for word_id in dictionary.token2id.values():
    word = dictionary.id2token[word_id]
    # 获取词在各主题下的概率分布
    topic_probs = lda_model.get_topic_terms(word_id, topn=lda_model.num_topics)
    # 整理成完整的主题概率数组(补全未返回的主题概率为0)
    full_probs = np.zeros(lda_model.num_topics)
    for topic_idx, prob in topic_probs:
        full_probs[topic_idx] = prob
    word_topic_dist.append(full_probs)
    # 统计词的总出现频次
    freq = sum(count for doc in corpus for (wid, count) in doc if wid == word_id)
    word_freq.append(freq)
    words.append(word)

# 2. 对词的主题分布做TSNE降维
tsne_model = TSNE(n_components=2, verbose=1, random_state=0, angle=.99, init='pca')
tsne_words = tsne_model.fit_transform(np.array(word_topic_dist))

# 3. 确定每个词的主导主题
dominant_topic = np.argmax(np.array(word_topic_dist), axis=1)

# 4. 准备Bokeh数据源
n_topics = lda_model.num_topics
mycolors = np.array([color for name, color in mcolors.TABLEAU_COLORS.items()[:n_topics]])

# 过滤掉低频词(可选,避免图表过于拥挤)
min_freq = 5
filtered_data = pd.DataFrame({
    'x': tsne_words[:,0],
    'y': tsne_words[:,1],
    'word': words,
    'freq': word_freq,
    'topic': dominant_topic,
    'color': mycolors[dominant_topic]
}).query(f'freq >= {min_freq}')

# 调整词大小:将词频映射到合适的字号范围(可根据需求调整系数)
filtered_data['font_size'] = filtered_data['freq'].apply(lambda x: min(8 + x*0.2, 20))  # 限制最大字号为20

source = ColumnDataSource(filtered_data)

# 5. 绘制可视化图表
output_notebook()
plot = figure(
    title="t-SNE Clustering of LDA Topics (Word Labels by Frequency)",
    plot_width=1500,
    plot_height=800,
    tools="pan,zoom_in,zoom_out,reset"
)

# 添加词标签
labels = LabelSet(
    x='x', y='y', text='word',
    text_color='color', text_font_size='font_size',
    source=source, x_offset=5, y_offset=5
)
plot.add_layout(labels)

# 添加主题图例(可选)
legend_source = ColumnDataSource(data=dict(
    topic_num=[str(i) for i in range(n_topics)],
    color=mycolors[:n_topics]
))
plot.circle(x=0, y=0, fill_color='color', line_color=None, size=10, legend_field='topic_num', source=legend_source)
plot.legend.title = "Topic ID"
plot.legend.location = "top_left"
plot.legend.label_text_font_size = "12pt"

show(plot)

关键修改说明

  • 词主题分布与词频统计:遍历词典中的每个词,获取其在LDA各主题下的概率分布,并统计该词在整个语料库中的总出现频次
  • TSNE降维对象切换:从原有的文档主题权重改为对词主题分布进行降维,让语义/主题相近的词在图中聚集
  • 词标签绘制:使用Bokeh的LabelSet替代散点,直接展示词文本
  • 词大小映射:将词频转换为字号,确保高频词显示更大,低频词更小(可通过调整font_size的计算系数优化显示效果)
  • 可选过滤:添加低频词过滤,避免图表因大量低频词变得杂乱

内容的提问来源于stack exchange,提问作者alex Maia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 11:09:53