基于textmineR的LDA模型:如何获取单篇文档的主题标签?
Hey there! I see you've already built an LDA model with textmineR and added topic labels using LabelTopics. Now to get the corresponding topic labels for the first 100 documents in nih_sample$ABSTRACT_TEXT, here are two practical approaches depending on your needs:
Approach 1: Get the dominant topic label (highest probability)
This assigns one primary topic to each document, based on the highest topic probability score in model$theta:
# Extract topic probabilities for the first 100 documents theta_top100 <- model$theta[1:100, ] # Find the index of the topic with the highest probability for each document dominant_topic_indices <- apply(theta_top100, 1, which.max) # Map these indices to your pre-defined topic labels document_dominant_labels <- model$labels[dominant_topic_indices] # Combine with original document metadata for easy reference document_topic_df <- data.frame( Application_ID = rownames(theta_top100), Abstract = nih_sample$ABSTRACT_TEXT[1:100], Dominant_Topic_Label = document_dominant_labels, stringsAsFactors = FALSE ) # Preview the first few results head(document_topic_df)
Approach 2: Get all relevant topic labels (probability > 0.05)
If you want to include every topic that meets your threshold (0.05, which you used for LabelTopics), this method collects and concatenates all matching labels for each document:
# Extract topic probabilities for the first 100 documents theta_top100 <- model$theta[1:100, ] # Find all topics with probability > 0.05 for each document relevant_topic_indices <- apply(theta_top100, 1, function(probs) which(probs > 0.05)) # Map indices to labels and join multiple labels with commas document_relevant_labels <- sapply(relevant_topic_indices, function(indices) { paste(model$labels[indices], collapse = ", ") }) # Combine with original document data document_topic_df <- data.frame( Application_ID = rownames(theta_top100), Abstract = nih_sample$ABSTRACT_TEXT[1:100], Relevant_Topic_Labels = document_relevant_labels, stringsAsFactors = FALSE ) # Preview the first few results head(document_topic_df)
Quick Notes:
model$thetais a matrix where rows match your documents (aligned with the order ofdtm) and columns match your 20 topics. Each value represents how likely a document belongs to that topic.- Your
model$labelsvector is indexed to match topics directly—somodel$labels[3]is the label for topic 3, making it easy to map indices to labels. - The row names of
theta_top100are theAPPLICATION_IDvalues from your original data, which ensures everything stays aligned correctly.
内容的提问来源于stack exchange,提问作者ch.elahe
相关产品推荐
相关产品推荐

