You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中如何按指定关键词Hirfar将文本拆分为多个独立章节片段

R语言基于指定关键词拆分文本的实现方法

方案1:直接对原文本做分割(推荐,保留原有格式)

不需要提前将文本拆分为单词,直接按关键词匹配拆分即可,会自动保留原有的标点、段落、换行等格式:

library(stringr)
# 按关键词"Hirfar"分割,匹配前后任意空白字符(含空格、换行)
split_chapters <- str_split(t, pattern = "\\s*Hirfar\\s*")[[1]]
# 过滤分割后产生的首段空值(因原文本开头就有关键词)
split_chapters <- split_chapters[nzchar(split_chapters)]

# 验证输出:共5个独立章节
length(split_chapters) 

# 输出各章节内容示例
for (i in 1:length(split_chapters)) {
  cat(paste0("第",i,"章:\n", split_chapters[i], "\n\n"))
}

方案2:基于已拆分的单词向量实现

如果需要沿用你已经拆分好的单词向量,可以通过定位关键词位置提取内容,注意该方法会丢失原有的标点、格式信息:

# 补全你之前的单词拆分代码
wrds <- str_split(t, pattern = boundary(type = "word"))[[1]]
# 定位所有"Hirfar"的索引位置
hirfar_index <- which(wrds == "Hirfar")
chapter_res <- list()

for (i in seq_along(hirfar_index)) {
  # 计算当前章节的单词起止范围
  start_pos <- hirfar_index[i] + 1
  end_pos <- ifelse(i == length(hirfar_index), length(wrds), hirfar_index[i+1] - 1)
  # 拼接单词为完整章节文本
  chapter_res[[i]] <- paste(wrds[start_pos:end_pos], collapse = " ")
}

# 查看拆分结果
chapter_res

内容的提问来源于stack exchange,提问作者Mohammad Farsadnia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 08:45:04