R语言求助:如何用正则表达式拆分文本字符串为前后两列?
问题描述
我有一个包含10列的DataFrame,其中一列是混杂数字、字符、标点及大写字母的长字符串。需要将该列拆分为before和includingandAfter两列:
includingandAfter:以任意4个连续字母开头,遇到标点时不停止读取before:提取正则匹配部分之前的所有内容
此前使用str_extract_all函数,将匹配固定词“DEAR”的正则替换为匹配4个连续字母的正则时出现报错,希望实现提取匹配部分之前的内容,以及从匹配部分开始往后的所有内容。
示例代码
# 构造测试数据 rant <- data.frame( reviews = c( "2022-01-22 DEAR diary, wish I could use SAS & b done in 10 min. with SUBSCR and Index", "2022-01-23 DEAR DIARY - I hope someone can help" ) ) # 原匹配固定词"DEAR"的代码 str_extract_all( rant, pattern = "(\\w+\\s){0,100} DEAR \\s?(\\w+\\s){0,100}" ) # 报错的匹配4个连续字母的代码 str_extract_all( rant, pattern = "(\\w+\\s){0,10} \\"\\s[A-Za-z]{4}\\" \\s?(\\w+\\s){0,10}" )
期望输出
拆分后得到如下两列:
| before | includingandAfter |
|---|---|
| 2022-01-22 | DEAR diary, wish I .... |
| 2022-01-23 | DEAR DIARY - I hope someo... |
解决方案
方法1:使用tidyr::separate拆分
利用零宽断言定位拆分点,直接将字符串拆分为两列:
library(tidyr) library(stringr) rant <- data.frame( reviews = c( "2022-01-22 DEAR diary, wish I could use SAS & b done in 10 min. with SUBSCR and Index", "2022-01-23 DEAR DIARY - I hope someone can help" ) ) rant_split <- rant %>% separate(reviews, into = c("before", "includingandAfter"), sep = "(?<=\\S)(?=[A-Za-z]{4})", # 匹配非空白字符后、4连续字母前的位置 extra = "merge", # 把剩余内容合并到第二列 fill = "right") %>% mutate(before = str_trim(before)) # 清理before列末尾空格 print(rant_split)
方法2:使用stringr::str_match提取分组
通过正则分组直接捕获前后两部分内容:
library(stringr) rant <- data.frame( reviews = c( "2022-01-22 DEAR diary, wish I could use SAS & b done in 10 min. with SUBSCR and Index", "2022-01-23 DEAR DIARY - I hope someone can help" ) ) # 正则分组:捕获匹配前内容 + 4连续字母开头的后续内容 pattern <- "^(.*?)\\s*([A-Za-z]{4}.*)$" matches <- str_match(rant$reviews, pattern) rant_split <- data.frame( before = matches[, 2], includingandAfter = matches[, 3] ) print(rant_split)
正则说明
^(.*?)\\s*([A-Za-z]{4}.*)$:^(.*?):非贪婪匹配开头到第一个4连续字母前的所有内容,避免过度匹配后续其他4字母组合\\s*:匹配4连续字母前的任意空格([A-Za-z]{4}.*):匹配以4个连续字母开头,包含所有后续内容(标点、数字等)的部分
(?<=\\S)(?=[A-Za-z]{4}):零宽断言,仅定位拆分位置,不消耗字符,确保拆分后内容完整
内容的提问来源于stack exchange,提问作者Sabor James
相关产品推荐
相关产品推荐

