You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:多字符分隔符分割字符串并保留分隔符的实现需求

R语言复杂字符串解析:按规则拆分并保留关键部分

需求说明

  • 条目拆分:长字符串中的每个条目以**可变长度数字+制表符(\t)**开头,拆分时需保留该条目编号前缀
  • 段落拆分:条目内的段落以「字符.字符」(无空格)为分隔特征,拆分时需保留前一段落末尾的句号

输入示例

input <- "1\tThis is a sentence. This is still part of the first paragraph.This is now the second paragraph.10\tThis is sentence number 1 of the tenth entry.This is the second sentence now. Still the second paragraph."

期望输出

[[1]]
[1] "1\tThis is a sentence. This is still part of the first paragraph."

[[2]]
[1] "This is now the second paragraph."

[[3]]
[1] "10\tThis is sentence number 1 of the tenth entry."

[[4]]
[1] "This is the second sentence now. Still the second paragraph."

解决方案

利用正则表达式的环视断言实现精准拆分,同时保留所需的分隔相关内容:

library(stringr)

input <- "1\tThis is a sentence. This is still part of the first paragraph.This is now the second paragraph.10\tThis is sentence number 1 of the tenth entry.This is the second sentence now. Still the second paragraph."

# 匹配两种拆分触发位置:句号后接条目编号(数字+制表符),或句号后接段落开头的大写字母
split_result <- str_split(input, "(?<=\\.)(?=\\d+\\t|[A-Z])")[[1]]

# 查看结果
split_result

正则逻辑解释

  • (?<=\.):正向回顾断言,确保拆分位置的前一个字符是句号,以此保留段落末尾的句号
  • (?=\d+\t|[A-Z]):正向先行断言,匹配两种触发拆分的后续内容:
    • \d+\t:一串数字加制表符(新条目的开头)
    • [A-Z]:大写字母(新段落的开头)

内容的提问来源于stack exchange,提问作者Aaron

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 22:25:13