如何在R中实现LDA主题时间变化分析(无需逐篇运行模型)
无需重复训练LDA追踪主题时间变化的方法
你已经基于月度聚合的演讲数据训练了全局LDA模型,完全可以直接利用现有模型输出分析主题的时间变化,不需要重复训练。核心思路是提取每个月度文档(即你聚合的每月演讲集合)的主题概率分布,再结合时间维度做趋势分析,具体步骤如下:
1. 提取月度文档的主题概率分布
训练好的speeches_lda_results已经包含了每个文档(对应一个月份)属于各个主题的概率(即gamma值),可以通过posterior()函数直接提取:
# 获取文档-主题概率矩阵 doc_topic_dist <- posterior(speeches_lda_results)$topics # 将矩阵的行名设置为对应的月份(和你之前split的月度列表顺序一致) month_labels <- names(PY_speeches_list) rownames(doc_topic_dist) <- month_labels # 转换为tidy格式,方便后续分析和可视化 library(tibble) library(tidyr) doc_topic_tidy <- doc_topic_dist %>% as_tibble(rownames = "date_month") %>% pivot_longer(cols = -date_month, names_to = "topic", values_to = "probability") %>% mutate( topic = as.integer(gsub("topic", "", topic)), # 把"topic1"转为数字1 date_month = lubridate::ym(date_month) # 转换为日期格式,方便时间排序 )
2. 可视化主题的时间趋势
有了tidy格式的数据,就可以轻松绘制每个主题随时间的概率变化:
单个主题的趋势
比如查看主题1的月度概率变化:
library(ggplot2) ggplot(doc_topic_tidy %>% filter(topic == 1), aes(x = date_month, y = probability)) + geom_line(color = "#2c3e50", linewidth = 1) + labs(title = "主题1的月度演讲占比变化", x = "月份", y = "主题概率") + theme_minimal()
所有主题的趋势(分面展示)
如果想同时观察多个主题的变化,用分面图:
ggplot(doc_topic_tidy, aes(x = date_month, y = probability)) + geom_line(color = "#3498db") + facet_wrap(~topic, scales = "free_y", ncol = 5) + # 按主题分面,每行5个 labs(title = "各主题的月度演讲占比变化", x = "月份", y = "主题概率") + theme_minimal() + theme(axis.text.x = element_text(angle = 45, hjust = 1))
3. 拓展:单篇演讲的主题时间变化
如果想细化到单篇演讲的主题分布,也不用重新训练模型,只需用训练好的LDA对单篇演讲做预测(注意预处理步骤要和训练时完全一致,且词汇表要匹配):
# 预处理单篇演讲(以PY_speeches中的第一条为例) single_speech <- Corpus(VectorSource(PY_speeches$text[1])) single_speech <- tm_map(single_speech, tolower) single_speech <- tm_map(single_speech, removePunctuation) single_speech <- tm_map(single_speech, function(x) removeWords(x, stopwords("english"))) single_speech <- tm_map(single_speech, stemDocument, language = "english") # 生成文档-词矩阵,必须使用训练时的词汇表(避免词汇不匹配) single_speech_dtm <- DocumentTermMatrix( single_speech, control = list( minWordLength=5, maxWordLength=12, dictionary = Terms(speeches_monthly_matrix) # 复用训练集的词汇表 ) ) # 预测该演讲的主题概率 single_speech_topic_prob <- predict(speeches_lda_results, newdata = single_speech_dtm)
之后把所有单篇演讲的主题概率和对应的date字段关联,就能绘制单篇演讲层面的主题时间变化曲线。
内容的提问来源于stack exchange,提问作者JF96
相关产品推荐
相关产品推荐

