纽约时报网页NLP场景下,如何使用自定义分隔符在tm语料库中拆分段落
解决在tm语料库中直接处理NYT段落分隔符的问题
看起来你卡在了不切换quanteda和tm就能处理纽约时报特殊段落分隔符的环节上,而且之前的gsub替换没生效大概率是因为那个•是Unicode特殊字符(U+2022),直接用普通字符串匹配可能出问题。下面给你一套完全基于tm的流程,不用来回转换格式,一步到位:
步骤1:爬取并预处理文本(先处理分隔符)
首先爬取网页后,先把•替换成换行符,再构建VCorpus,这样后续拆分段落更顺畅:
library(rvest) library(tm) library(stringr) # 爬取网页内容 website <- read_html("https://www.nytimes.com/2017/01/03/briefing/asia-australia-briefing.html") text <- website %>% html_nodes("p") %>% html_text() %>% as.character() # 关键:替换Unicode项目符号•为换行符(用Unicode转义确保匹配,避免编码问题) text_clean <- str_replace_all(text, "\u2022", "\n") # 合并文本后直接构建tm的VCorpus text_collapsed <- str_c(text_clean, collapse = "\n") corpus <- VCorpus(VectorSource(text_collapsed))
步骤2:在tm语料库中拆分段落并过滤目标内容
接下来用tm的tm_map结合自定义函数拆分段落,然后筛选包含目标关键词的段落:
# 自定义函数:将文档按换行符拆分为段落,过滤空段落 split_into_paragraphs <- function(doc) { paragraphs <- str_split(content(doc), "\n+")[[1]] paragraphs <- paragraphs[paragraphs != ""] lapply(paragraphs, PlainTextDocument) } # 拆分语料库为段落级文档并扁平化 corpus_paragraphs <- tm_map(corpus, split_into_paragraphs) corpus_paragraphs <- Corpus(VectorSource(sapply(corpus_paragraphs, content))) # 转小写并筛选包含目标短语的段落 corpus_paragraphs <- tm_map(corpus_paragraphs, content_transformer(tolower)) target_phrase <- "pull out of peace talks" corpus_filtered <- corpus_paragraphs[tm_map(corpus_paragraphs, content_transformer(str_detect), target_phrase)]
步骤3:构建DTM并统计词频
最后直接在tm里构建文档-词矩阵,统计高频词:
# 构建DTM(可按需添加停用词、去标点等预处理规则) dtm <- DocumentTermMatrix(corpus_filtered, control = list( stopwords = TRUE, removePunctuation = TRUE, stripWhitespace = TRUE )) # 查看出现至少1次的词 findFreqTerms(dtm, 1)
为什么之前的gsub没生效?
那个•是Unicode项目符号(U+2022),如果你的R环境编码设置不对,直接写gsub("•", "\n", text)可能匹配不到。用"\u2022"或者直接从网页复制粘贴该符号到代码里,就能确保匹配成功。
这样整个流程都在tm生态里完成,不用再转quanteda或者数据框,操作简洁多了。
内容的提问来源于stack exchange,提问作者SLE
相关产品推荐
相关产品推荐

