You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中实现LDA主题时间变化分析(无需逐篇运行模型)

无需重复训练LDA追踪主题时间变化的方法

你已经基于月度聚合的演讲数据训练了全局LDA模型,完全可以直接利用现有模型输出分析主题的时间变化,不需要重复训练。核心思路是提取每个月度文档(即你聚合的每月演讲集合)的主题概率分布,再结合时间维度做趋势分析,具体步骤如下:

1. 提取月度文档的主题概率分布

训练好的speeches_lda_results已经包含了每个文档(对应一个月份)属于各个主题的概率(即gamma值),可以通过posterior()函数直接提取:

# 获取文档-主题概率矩阵
doc_topic_dist <- posterior(speeches_lda_results)$topics

# 将矩阵的行名设置为对应的月份(和你之前split的月度列表顺序一致)
month_labels <- names(PY_speeches_list)
rownames(doc_topic_dist) <- month_labels

# 转换为tidy格式,方便后续分析和可视化
library(tibble)
library(tidyr)
doc_topic_tidy <- doc_topic_dist %>%
  as_tibble(rownames = "date_month") %>%
  pivot_longer(cols = -date_month, names_to = "topic", values_to = "probability") %>%
  mutate(
    topic = as.integer(gsub("topic", "", topic)), # 把"topic1"转为数字1
    date_month = lubridate::ym(date_month) # 转换为日期格式,方便时间排序
  )

2. 可视化主题的时间趋势

有了tidy格式的数据,就可以轻松绘制每个主题随时间的概率变化:

单个主题的趋势

比如查看主题1的月度概率变化:

library(ggplot2)
ggplot(doc_topic_tidy %>% filter(topic == 1), aes(x = date_month, y = probability)) +
  geom_line(color = "#2c3e50", linewidth = 1) +
  labs(title = "主题1的月度演讲占比变化", x = "月份", y = "主题概率") +
  theme_minimal()

所有主题的趋势(分面展示)

如果想同时观察多个主题的变化,用分面图:

ggplot(doc_topic_tidy, aes(x = date_month, y = probability)) +
  geom_line(color = "#3498db") +
  facet_wrap(~topic, scales = "free_y", ncol = 5) + # 按主题分面,每行5个
  labs(title = "各主题的月度演讲占比变化", x = "月份", y = "主题概率") +
  theme_minimal() +
  theme(axis.text.x = element_text(angle = 45, hjust = 1))

3. 拓展:单篇演讲的主题时间变化

如果想细化到单篇演讲的主题分布,也不用重新训练模型,只需用训练好的LDA对单篇演讲做预测(注意预处理步骤要和训练时完全一致,且词汇表要匹配):

# 预处理单篇演讲(以PY_speeches中的第一条为例)
single_speech <- Corpus(VectorSource(PY_speeches$text[1]))
single_speech <- tm_map(single_speech, tolower)
single_speech <- tm_map(single_speech, removePunctuation)
single_speech <- tm_map(single_speech, function(x) removeWords(x, stopwords("english")))
single_speech <- tm_map(single_speech, stemDocument, language = "english")

# 生成文档-词矩阵,必须使用训练时的词汇表(避免词汇不匹配)
single_speech_dtm <- DocumentTermMatrix(
  single_speech, 
  control = list(
    minWordLength=5, 
    maxWordLength=12, 
    dictionary = Terms(speeches_monthly_matrix) # 复用训练集的词汇表
  )
)

# 预测该演讲的主题概率
single_speech_topic_prob <- predict(speeches_lda_results, newdata = single_speech_dtm)

之后把所有单篇演讲的主题概率和对应的date字段关联,就能绘制单篇演讲层面的主题时间变化曲线。

内容的提问来源于stack exchange,提问作者JF96

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 12:45:06