You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

纽约时报网页NLP场景下,如何使用自定义分隔符在tm语料库中拆分段落

解决在tm语料库中直接处理NYT段落分隔符的问题

看起来你卡在了不切换quanteda和tm就能处理纽约时报特殊段落分隔符的环节上,而且之前的gsub替换没生效大概率是因为那个•是Unicode特殊字符(U+2022),直接用普通字符串匹配可能出问题。下面给你一套完全基于tm的流程,不用来回转换格式,一步到位:

步骤1:爬取并预处理文本(先处理分隔符)

首先爬取网页后,先把•替换成换行符,再构建VCorpus,这样后续拆分段落更顺畅:

library(rvest)
library(tm)
library(stringr)

# 爬取网页内容
website <- read_html("https://www.nytimes.com/2017/01/03/briefing/asia-australia-briefing.html")
text <- website %>% html_nodes("p") %>% html_text() %>% as.character()

# 关键:替换Unicode项目符号•为换行符(用Unicode转义确保匹配,避免编码问题)
text_clean <- str_replace_all(text, "\u2022", "\n")

# 合并文本后直接构建tm的VCorpus
text_collapsed <- str_c(text_clean, collapse = "\n")
corpus <- VCorpus(VectorSource(text_collapsed))

步骤2:在tm语料库中拆分段落并过滤目标内容

接下来用tm的tm_map结合自定义函数拆分段落,然后筛选包含目标关键词的段落:

# 自定义函数:将文档按换行符拆分为段落,过滤空段落
split_into_paragraphs <- function(doc) {
  paragraphs <- str_split(content(doc), "\n+")[[1]]
  paragraphs <- paragraphs[paragraphs != ""]
  lapply(paragraphs, PlainTextDocument)
}

# 拆分语料库为段落级文档并扁平化
corpus_paragraphs <- tm_map(corpus, split_into_paragraphs)
corpus_paragraphs <- Corpus(VectorSource(sapply(corpus_paragraphs, content)))

# 转小写并筛选包含目标短语的段落
corpus_paragraphs <- tm_map(corpus_paragraphs, content_transformer(tolower))
target_phrase <- "pull out of peace talks"
corpus_filtered <- corpus_paragraphs[tm_map(corpus_paragraphs, content_transformer(str_detect), target_phrase)]

步骤3:构建DTM并统计词频

最后直接在tm里构建文档-词矩阵,统计高频词:

# 构建DTM(可按需添加停用词、去标点等预处理规则)
dtm <- DocumentTermMatrix(corpus_filtered, control = list(
  stopwords = TRUE,
  removePunctuation = TRUE,
  stripWhitespace = TRUE
))

# 查看出现至少1次的词
findFreqTerms(dtm, 1)

为什么之前的gsub没生效?

那个•是Unicode项目符号(U+2022),如果你的R环境编码设置不对,直接写gsub("•", "\n", text)可能匹配不到。用"\u2022"或者直接从网页复制粘贴该符号到代码里,就能确保匹配成功。

这样整个流程都在tm生态里完成,不用再转quanteda或者数据框,操作简洁多了。

内容的提问来源于stack exchange,提问作者SLE

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 13:17:46