You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

主题建模中如何从tmResult$terms提取术语与概率并生成主题词云?

为LDA模型的8个主题生成独立词云的解决方案

问题描述

我想给LDA模型中的8个主题分别生成独立词云。已经提取了每个主题的40个高频词及对应出现概率,得到了长度为320的top_words_vector对象,但无法从中拆分出术语和概率值。以下是我的代码:

textdata <- base::readRDS(url("https://slcladal.github.io/data/sotu_paragraphs.rda", "rb"))

english_stopwords <- readLines("https://slcladal.github.io/resources/stopwords_en.txt", encoding = "UTF-8")

corpus <- Corpus(DataframeSource(textdata))

processedCorpus <- tm_map(corpus, content_transformer(tolower))
processedCorpus <- tm_map(processedCorpus, removeWords, english_stopwords)
processedCorpus <- tm_map(processedCorpus, removePunctuation, preserve_intra_word_dashes = TRUE)
processedCorpus <- tm_map(processedCorpus, removeNumbers)
processedCorpus <- tm_map(processedCorpus, stemDocument, language = "en")
processedCorpus <- tm_map(processedCorpus, stripWhitespace)

minimumFrequency <- 5
DTM <- DocumentTermMatrix(processedCorpus, control = list(bounds = list(global = c(minimumFrequency, Inf))))
sel_idx <- slam::row_sums(DTM) > 0
DTM <- DTM[sel_idx, ]
textdata <- textdata[sel_idx, ]

K <- 8
set.seed(9161)
# compute the LDA model, inference via 100 iterations of Gibbs sampling
topicModel <- LDA(DTM, K, method="Gibbs", control=list(iter = 100, verbose = 25))
tmResult <- topicmodels::posterior(topicModel)
tmResult$terms 

top_words_vector = c() # an empty container for 320 length, top#40 words across 8 topics
for(i in 1:8){
  top_words_vector = c(top_words_vector,sort(tmResult$terms[i,], decreasing=TRUE)[1:40])
}

top_words_vector

wordcloud()函数需要分别传入术语和概率参数,我正尝试从top_words_vector中提取这两类数据:

mycolors <- brewer.pal(8, "Dark2")
wordcloud(c("apple", "banana"), c(0.8,0.2), random.order = TRUE, color = mycolors)

解决方案

问题核心是你将所有主题的概率值合并成了一维向量,丢失了对应的术语名称。需要保留术语与概率的对应关系,再批量生成词云。

1. 重新存储主题的高频词与概率

用列表存储每个主题的术语-概率对应数据,替代原有的一维向量:

# 创建列表存储每个主题的Top40词及概率
top_words_list <- list()
for(i in 1:K){
  # 对当前主题的词概率降序排序,取前40个
  sorted_terms <- sort(tmResult$terms[i,], decreasing = TRUE)[1:40]
  # 存储为数据框,保留术语和概率
  top_words_list[[i]] <- data.frame(
    term = names(sorted_terms),
    prob = as.numeric(sorted_terms),
    stringsAsFactors = FALSE
  )
}

2. 循环生成每个主题的词云

遍历列表中的每个主题数据,调用wordcloud()生成独立词云:

library(wordcloud)
library(RColorBrewer)

mycolors <- brewer.pal(8, "Dark2")

# 逐个生成主题词云
for(i in 1:K){
  current_data <- top_words_list[[i]]
  # 生成词云,添加主题标题区分
  wordcloud(words = current_data$term, 
            freq = current_data$prob,
            random.order = FALSE,  # 按频率排序,高频词更显眼
            colors = mycolors,
            main = paste("主题", i, "词云"))
}

关键说明

  • 用列表+数据框的结构,能完整保留每个主题的术语与概率对应关系,避免信息丢失。
  • random.order = FALSE让高频词显示在词云中心区域,提升可读性。
  • 每个词云会自动弹出窗口(或在RStudio的Plots面板中切换查看),清晰区分8个主题。

内容的提问来源于stack exchange,提问作者NoaMi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 23:19:51