如何结合Gensim LDA与TSNE绘制带词频权重的词标签聚类图?
改造LDA主题可视化:词标签替代散点并按词频调整大小
以下是修改后的完整代码,实现用词标签替代原散点图,并且词的大小与该词在语料库中的出现频次成正比:
from gensim.models import LdaModel from gensim.corpora.dictionary import Dictionary from sklearn.manifold import TSNE from bokeh.plotting import figure, output_notebook, show from bokeh.models import ColumnDataSource, LabelSet import pandas as pd import numpy as np import matplotlib.colors as mcolors # --- 你的原有LDA模型代码(保留即可)--- # dictionary = Dictionary(all_texts) # corpus = [dictionary.doc2bow(text) for text in all_texts] # lda_model = LdaModel(corpus=corpus, id2word=dictionary, num_topics=10, alpha="auto",per_word_topics=True,random_state=42) # 1. 获取每个词的主题分布和词频 word_topic_dist = [] word_freq = [] words = [] for word_id in dictionary.token2id.values(): word = dictionary.id2token[word_id] # 获取词在各主题下的概率分布 topic_probs = lda_model.get_topic_terms(word_id, topn=lda_model.num_topics) # 整理成完整的主题概率数组(补全未返回的主题概率为0) full_probs = np.zeros(lda_model.num_topics) for topic_idx, prob in topic_probs: full_probs[topic_idx] = prob word_topic_dist.append(full_probs) # 统计词的总出现频次 freq = sum(count for doc in corpus for (wid, count) in doc if wid == word_id) word_freq.append(freq) words.append(word) # 2. 对词的主题分布做TSNE降维 tsne_model = TSNE(n_components=2, verbose=1, random_state=0, angle=.99, init='pca') tsne_words = tsne_model.fit_transform(np.array(word_topic_dist)) # 3. 确定每个词的主导主题 dominant_topic = np.argmax(np.array(word_topic_dist), axis=1) # 4. 准备Bokeh数据源 n_topics = lda_model.num_topics mycolors = np.array([color for name, color in mcolors.TABLEAU_COLORS.items()[:n_topics]]) # 过滤掉低频词(可选,避免图表过于拥挤) min_freq = 5 filtered_data = pd.DataFrame({ 'x': tsne_words[:,0], 'y': tsne_words[:,1], 'word': words, 'freq': word_freq, 'topic': dominant_topic, 'color': mycolors[dominant_topic] }).query(f'freq >= {min_freq}') # 调整词大小:将词频映射到合适的字号范围(可根据需求调整系数) filtered_data['font_size'] = filtered_data['freq'].apply(lambda x: min(8 + x*0.2, 20)) # 限制最大字号为20 source = ColumnDataSource(filtered_data) # 5. 绘制可视化图表 output_notebook() plot = figure( title="t-SNE Clustering of LDA Topics (Word Labels by Frequency)", plot_width=1500, plot_height=800, tools="pan,zoom_in,zoom_out,reset" ) # 添加词标签 labels = LabelSet( x='x', y='y', text='word', text_color='color', text_font_size='font_size', source=source, x_offset=5, y_offset=5 ) plot.add_layout(labels) # 添加主题图例(可选) legend_source = ColumnDataSource(data=dict( topic_num=[str(i) for i in range(n_topics)], color=mycolors[:n_topics] )) plot.circle(x=0, y=0, fill_color='color', line_color=None, size=10, legend_field='topic_num', source=legend_source) plot.legend.title = "Topic ID" plot.legend.location = "top_left" plot.legend.label_text_font_size = "12pt" show(plot)
关键修改说明
- 词主题分布与词频统计:遍历词典中的每个词,获取其在LDA各主题下的概率分布,并统计该词在整个语料库中的总出现频次
- TSNE降维对象切换:从原有的文档主题权重改为对词主题分布进行降维,让语义/主题相近的词在图中聚集
- 词标签绘制:使用Bokeh的
LabelSet替代散点,直接展示词文本 - 词大小映射:将词频转换为字号,确保高频词显示更大,低频词更小(可通过调整
font_size的计算系数优化显示效果) - 可选过滤:添加低频词过滤,避免图表因大量低频词变得杂乱
内容的提问来源于stack exchange,提问作者alex Maia
相关产品推荐
相关产品推荐

