使用R语言tm包构建语料库时遇错误求助
问题排查:tm包构建语料库时的DataframeSource错误
错误原因
你构建语料库时调用了DataframeSource(df),但原始的df仅包含text和dates两个字段,缺少DataframeSource要求的必填字段doc_id,因此触发字段匹配失败的错误。而你实际已经准备好符合要求的df.tm(包含doc_id和text),只是最后一行代码传错了数据集。
解决步骤
- 将最后一行的
df替换为你预先处理好的df.tm - 可选验证:执行
names(df.tm)确认输出为[1] "doc_id" "text",确保列名完全符合要求
修正后的关键代码片段
# 从评论文本创建语料库(修正后) corpus <- VCorpus(DataframeSource(df.tm))
完整修正代码
library(tidyverse) library(rvest) ########################################## # WEB SCRAPING FROM SCHOLARLYKITCHEN.COM # ########################################## # 创建循环迭代获取页面链接(测试用小范围循环) output <- character() for (i in 1:2) { article.links <- paste0("https://scholarlykitchen.sspnet.org/archives/page/", i ,"/") %>% read_html() %>% html_nodes(".list-article__title") %>% html_nodes("a") %>% html_attr("href") output <- c(output, article.links) } # 获取评论内容 get.comments <- function(output) { article.page <- read_html(output) article.comments <- article.page %>% html_nodes(".comment") %>% html_text() %>% trimws(which = "both") return(article.comments) } text <- sapply(output, FUN = get.comments, USE.NAMES = FALSE) # 获取评论日期 get.dates <- function(output) { article.page <- read_html(output) article.comments <- article.page %>% html_nodes(".comment__meta__date") %>% html_text() %>% trimws(which = "both") return(article.comments) } dates <- sapply(output, FUN = get.dates, USE.NAMES = FALSE) # 构建分析用数据集 df <- tibble( text = unlist(text, recursive = TRUE), dates = unlist(dates, recursive = TRUE) ) # 日期格式转换 df$dates <- as.character(gsub(",","",df$dates)) df$dates <- as.Date(df$dates, "%B%d%Y") ################### # TOPIC MODELLING # ################### library(tm) library(topicmodels) # 构建主题建模专用数据集(包含doc_id和text) df.tm <- df[-2] df.tm$doc_id <- row.names(df) df.tm <- df.tm[c(2,1)] # 从评论文本创建语料库(已修正) corpus <- VCorpus(DataframeSource(df.tm))
内容的提问来源于stack exchange,提问作者I_like_insights
相关产品推荐
相关产品推荐

