You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言求助:如何用正则表达式拆分文本字符串为前后两列?

问题描述

我有一个包含10列的DataFrame,其中一列是混杂数字、字符、标点及大写字母的长字符串。需要将该列拆分为before和includingandAfter两列:

  • includingandAfter:以任意4个连续字母开头,遇到标点时不停止读取
  • before:提取正则匹配部分之前的所有内容

此前使用str_extract_all函数,将匹配固定词“DEAR”的正则替换为匹配4个连续字母的正则时出现报错,希望实现提取匹配部分之前的内容,以及从匹配部分开始往后的所有内容。

示例代码
# 构造测试数据
rant <- data.frame(
  reviews = c(
    "2022-01-22 DEAR diary, wish I could use SAS & b done in 10 min. with SUBSCR and Index",
    "2022-01-23 DEAR DIARY - I hope someone can help"
  )
)

# 原匹配固定词"DEAR"的代码
str_extract_all(
  rant,
  pattern = "(\\w+\\s){0,100} DEAR \\s?(\\w+\\s){0,100}"
)

# 报错的匹配4个连续字母的代码
str_extract_all(
  rant,
  pattern = "(\\w+\\s){0,10} \\"\\s[A-Za-z]{4}\\"  \\s?(\\w+\\s){0,10}"
)
期望输出

拆分后得到如下两列:

beforeincludingandAfter
2022-01-22DEAR diary, wish I ....
2022-01-23DEAR DIARY - I hope someo...
解决方案

方法1:使用tidyr::separate拆分

利用零宽断言定位拆分点,直接将字符串拆分为两列:

library(tidyr)
library(stringr)

rant <- data.frame(
  reviews = c(
    "2022-01-22 DEAR diary, wish I could use SAS & b done in 10 min. with SUBSCR and Index",
    "2022-01-23 DEAR DIARY - I hope someone can help"
  )
)

rant_split <- rant %>%
  separate(reviews, 
           into = c("before", "includingandAfter"), 
           sep = "(?<=\\S)(?=[A-Za-z]{4})", # 匹配非空白字符后、4连续字母前的位置
           extra = "merge", # 把剩余内容合并到第二列
           fill = "right") %>%
  mutate(before = str_trim(before)) # 清理before列末尾空格

print(rant_split)

方法2:使用stringr::str_match提取分组

通过正则分组直接捕获前后两部分内容:

library(stringr)

rant <- data.frame(
  reviews = c(
    "2022-01-22 DEAR diary, wish I could use SAS & b done in 10 min. with SUBSCR and Index",
    "2022-01-23 DEAR DIARY - I hope someone can help"
  )
)

# 正则分组:捕获匹配前内容 + 4连续字母开头的后续内容
pattern <- "^(.*?)\\s*([A-Za-z]{4}.*)$"
matches <- str_match(rant$reviews, pattern)

rant_split <- data.frame(
  before = matches[, 2],
  includingandAfter = matches[, 3]
)

print(rant_split)

正则说明

  • ^(.*?)\\s*([A-Za-z]{4}.*)$:
    • ^(.*?):非贪婪匹配开头到第一个4连续字母前的所有内容,避免过度匹配后续其他4字母组合
    • \\s*:匹配4连续字母前的任意空格
    • ([A-Za-z]{4}.*):匹配以4个连续字母开头,包含所有后续内容(标点、数字等)的部分
  • (?<=\\S)(?=[A-Za-z]{4}):零宽断言,仅定位拆分位置,不消耗字符,确保拆分后内容完整

内容的提问来源于stack exchange,提问作者Sabor James

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 16:46:30