You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于textmineR的LDA模型:如何获取单篇文档的主题标签?

Hey there! I see you've already built an LDA model with textmineR and added topic labels using LabelTopics. Now to get the corresponding topic labels for the first 100 documents in nih_sample$ABSTRACT_TEXT, here are two practical approaches depending on your needs:

Approach 1: Get the dominant topic label (highest probability)

This assigns one primary topic to each document, based on the highest topic probability score in model$theta:

# Extract topic probabilities for the first 100 documents
theta_top100 <- model$theta[1:100, ]

# Find the index of the topic with the highest probability for each document
dominant_topic_indices <- apply(theta_top100, 1, which.max)

# Map these indices to your pre-defined topic labels
document_dominant_labels <- model$labels[dominant_topic_indices]

# Combine with original document metadata for easy reference
document_topic_df <- data.frame(
  Application_ID = rownames(theta_top100),
  Abstract = nih_sample$ABSTRACT_TEXT[1:100],
  Dominant_Topic_Label = document_dominant_labels,
  stringsAsFactors = FALSE
)

# Preview the first few results
head(document_topic_df)

Approach 2: Get all relevant topic labels (probability > 0.05)

If you want to include every topic that meets your threshold (0.05, which you used for LabelTopics), this method collects and concatenates all matching labels for each document:

# Extract topic probabilities for the first 100 documents
theta_top100 <- model$theta[1:100, ]

# Find all topics with probability > 0.05 for each document
relevant_topic_indices <- apply(theta_top100, 1, function(probs) which(probs > 0.05))

# Map indices to labels and join multiple labels with commas
document_relevant_labels <- sapply(relevant_topic_indices, function(indices) {
  paste(model$labels[indices], collapse = ", ")
})

# Combine with original document data
document_topic_df <- data.frame(
  Application_ID = rownames(theta_top100),
  Abstract = nih_sample$ABSTRACT_TEXT[1:100],
  Relevant_Topic_Labels = document_relevant_labels,
  stringsAsFactors = FALSE
)

# Preview the first few results
head(document_topic_df)

Quick Notes:

  • model$theta is a matrix where rows match your documents (aligned with the order of dtm) and columns match your 20 topics. Each value represents how likely a document belongs to that topic.
  • Your model$labels vector is indexed to match topics directly—so model$labels[3] is the label for topic 3, making it easy to map indices to labels.
  • The row names of theta_top100 are the APPLICATION_ID values from your original data, which ensures everything stays aligned correctly.

内容的提问来源于stack exchange,提问作者ch.elahe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:32:19