You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用quanteda的corpus_segment函数分割管道分隔的演讲文本文件?

你现在的问题核心在于原始文本格式太乱——所有演讲ID和内容都挤在一起,没有清晰的分隔,直接用corpus_segment肯定没法正确识别。别慌,咱们分两步来搞定:先把文本拆成规范的ID-内容对,再导入quanteda生成独立文档。


步骤1:清理并解析原始文本

首先得把那串连在一起的内容拆成每个演讲ID对应一段文本。你的示例里,演讲内容结尾和下一个ID粘在了一起(比如This is the first speech.1140000002),咱们用正则表达式精准拆分:

先把文件内容读进R:

raw_text <- readLines("your_speeches.txt", warn = FALSE)

然后处理这段文本:

# 去掉开头多余的"Speech_id|speech"前缀
cleaned_text <- sub("Speech_id\\|speech", "", raw_text)

# 用正则匹配10位ID+|+内容(直到下一个ID或结尾)
matches <- gregexpr("(\\d{10})\\|(.*?)(?=\\d{10}\\||$)", cleaned_text, perl = TRUE)
speech_data <- do.call(rbind, regmatches(cleaned_text, matches))

# 转成结构化数据框
speech_df <- as.data.frame(speech_data, stringsAsFactors = FALSE)
colnames(speech_df) <- c("speech_id", "text")

这时候speech_df里就有清晰的两列:speech_id是每个演讲的唯一ID,text是对应的演讲内容,每行对应一个独立演讲。


步骤2:用quanteda生成独立文档

现在数据规范了,直接创建语料库,每个演讲自动成为一个独立文档:

library(quanteda)

# 从数据框创建语料库,指定文本列和文档ID列
speech_corpus <- corpus(speech_df, text_field = "text", docid_field = "speech_id")

# 查看结果,确认每个文档对应一个演讲
summary(speech_corpus)

如果你确实需要用corpus_segment拆分更细的单元(比如句子),现在有了规范的语料库就好办了:

# 把每个演讲拆分成句子级的子文档
segmented_corpus <- corpus_segment(speech_corpus, pattern = "\\.", valuetype = "regex")

为啥之前没成功?

你之前尝试失败大概率是因为原始文本格式不规范——ID和内容没有明确分隔,corpus_segment需要清晰的拆分标记才能工作。先把数据整理成结构化格式,quanteda的工具才能发挥作用~

内容的提问来源于stack exchange,提问作者Justin W.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:15:59