如何将STM与LDA主题模型匹配原数据集并重命名主题?
为STM和LDA模型分配文档主题并重命名主题
STM模型:生成文档-主题对应数据框
STM模型的theta矩阵存储了每个文档对应各主题的概率值(行=文档,列=主题)。通过以下步骤可以生成包含文档ID和对应主题的数据框:
方法1:提取主导主题(概率最高的主题)
# 从dfm中获取文档ID doc_ids <- docnames(dfm_to_model) # 提取文档-主题概率矩阵 topic_probs <- stm_topic$theta # 找出每个文档概率最高的主题编号 dominant_topic <- apply(topic_probs, 1, which.max) # 生成最终数据框 stm_doc_topic_df <- data.frame( docid = doc_ids, dominant_topic = dominant_topic, stringsAsFactors = FALSE ) # 查看前几行结果 head(stm_doc_topic_df)
方法2:保留所有主题的概率值
如果需要完整的文档-主题概率分布,可以直接合并文档ID和概率矩阵:
stm_full_topic_df <- cbind(docid = doc_ids, as.data.frame(topic_probs)) head(stm_full_topic_df)
LDA模型(quanteda的textmodel_lda):生成文档-主题对应数据框
quanteda的textmodel_lda对象中,theta矩阵是主题×文档的结构,需要先转置为文档×主题格式,再处理:
# 获取文档ID doc_ids <- docnames(dfm_to_model) # 转置theta矩阵为文档×主题格式 lda_topic_probs <- t(tmod_lda$theta) # 提取每个文档的主导主题 lda_dominant_topic <- apply(lda_topic_probs, 1, which.max) # 生成数据框 lda_doc_topic_df <- data.frame( docid = doc_ids, dominant_topic = lda_dominant_topic, stringsAsFactors = FALSE ) # 查看结果 head(lda_doc_topic_df)
为主题重命名
无论是STM还是LDA模型,都可以通过映射关系将默认的主题编号替换为自定义名称:
STM模型主题重命名
# 定义主题名称映射(根据实际主题内容调整) topic_name_map <- c( "1" = "Leadership", "2" = "National Unity", "3" = "Policy & Governance", "4" = "Foreign Affairs", "5" = "Civil Rights" ) # 方法1:在数据框中替换主题编号为自定义名称 stm_doc_topic_df$topic_name <- recode(stm_doc_topic_df$dominant_topic, !!!topic_name_map) # 方法2:为STM模型添加主题标签(可视化时自动显示) stm_topic$topic.names <- unname(topic_name_map) # 重新绘图会显示自定义主题名 plot(stm_topic, type="summary")
LDA模型主题重命名
# 使用同样的主题名称映射 topic_name_map <- c( "1" = "Leadership", "2" = "National Unity", "3" = "Policy & Governance", "4" = "Foreign Affairs", "5" = "Civil Rights" ) # 方法1:修改数据框中的主题名称 lda_doc_topic_df$topic_name <- recode(lda_doc_topic_df$dominant_topic, !!!topic_name_map) # 方法2:修改LDA模型的主题标签(提取术语时显示自定义名称) names(tmod_lda$topics) <- unname(topic_name_map) # 提取术语时会显示自定义主题名 seededlda::terms(tmod_lda, 10)
内容的提问来源于stack exchange,提问作者R_Student
相关产品推荐
相关产品推荐

