如何为主题建模结果分配列名?解决赋值长度不匹配错误
主题建模代码报错:'names' attribute [3] must be the same length as the vector [1]
错误原因
你创建的document_topics是仅含1列的数据框(存储每个文档对应的最可能主题编号),但你试图给它设置3个列名——列名数量和数据框的列数不匹配,这就是报错的核心原因。
修正方案
根据你的需求,分两种场景提供解决方案:
场景1:给每个文档的主题编号替换成自定义标签
如果只是想把文档对应的主题数字(1/2/3)换成你定义的主题名称,修改代码中document_topics相关的部分:
topics <- topics(lda_model) # 将主题编号映射为自定义标签 document_topics <- as.data.frame(Assigned_Topic = factor(topics, levels = 1:num_topics, labels = topic_labels)) print(document_topics)
场景2:展示每个文档在3个主题上的概率分布
如果想显示每个文档在所有主题上的概率(此时数据框会有3列,和你的3个标签匹配),改用posterior()提取主题概率矩阵:
# 提取每个文档的主题概率分布 topic_probs <- as.data.frame(posterior(lda_model)$topics) # 设置列名为自定义主题标签 colnames(topic_probs) <- topic_labels print(topic_probs)
完整修正代码(场景1示例)
install.packages("tm") install.packages("topicmodels") library(tm) library(topicmodels) docs <- Corpus(VectorSource(c( "This is the first document about topic modeling.", "Topic modeling is a popular technique in text analysis.", "LDA is a common algorithm for topic modeling.", "Text mining is an interesting field in data science." ))) docs <- tm_map(docs, content_transformer(tolower)) # 转小写 docs <- tm_map(docs, removePunctuation) # 移除标点 docs <- tm_map(docs, removeNumbers) # 移除数字 docs <- tm_map(docs, removeWords, stopwords("english")) # 移除停用词 docs <- tm_map(docs, stripWhitespace) dtm <- DocumentTermMatrix(docs) num_topics <- 3 lda_model <- LDA(dtm, k = num_topics) topic_labels <- c("Topic 1: Introduction to Topic Modeling", "Topic 2: Techniques in Text Analysis", "Topic 3: LDA and Data Science") terms(lda_model, 10) # 修正后的主题映射代码 topics <- topics(lda_model) document_topics <- as.data.frame(Assigned_Topic = factor(topics, levels = 1:num_topics, labels = topic_labels)) print(document_topics)
内容的提问来源于stack exchange,提问作者Junaid
相关产品推荐
相关产品推荐

