You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中提取多层嵌套括号内的code IS (...)字符串

提取带嵌套括号的code IS (...)实例

原始目标字符串

mystring <- c("code IS (k(384333)
   AND parse = TURE 
 ) 

              code IS (
 FROM (43343344)
 ) some information code IS 

              code IS (  ( 
 (data)(23423422 
)) ) ) and more information)")

当前问题:无法匹配嵌套括号

用普通正则提取时,只会匹配到第一个闭合括号就终止,结果不符合预期:

library(stringr)
> str_extract_all(pattern = 'code IS \\([\\s\\S]+?\\)', mystring)
[[1]]
[1] "code IS (k(384333)"          "code IS (
 FROM (43343344)" "code IS (  ( 
 (data)"    

期望输出

[[1]]
[1] "code IS (k(384333)
   AND parse = TURE 
 )"          "code IS (
 FROM (43343344)
 )" "code IS (  ( 
 (data)(23423422 
)) )" 

适配平衡括号正则的报错问题

尝试用递归正则匹配平衡括号时,因R的字符串转义和引擎支持问题报错:

> str_extract_all(pattern = 'code IS \((?:[^)(]+|(?R))*+\)', mystring)
Error: '\(' is an unrecognized escape in character string starting "'code IS \(" 

解决方法

R中要注意两点:一是字符串里的正则元字符需要双重转义,二是stringr默认用ICU引擎,不支持递归正则,得指定用PCRE引擎(加perl=TRUE参数)。

正确的代码如下:

library(stringr)
str_extract_all(mystring, pattern = 'code IS \\((?:[^()]|(?R))*+\\)', perl = TRUE)

正则说明

  • \\(/\\):R字符串中括号需要转两次,才能被正则引擎识别为括号元字符
  • (?:[^()]|(?R))*+:非捕获组,要么匹配非括号的任意字符,要么递归匹配整个括号表达式((?R)是PCRE的递归语法,专门用来处理嵌套平衡结构)
  • perl=TRUE:启用PCRE引擎,支持递归匹配逻辑

运行这段代码就能得到符合预期的嵌套括号匹配结果。


内容的提问来源于stack exchange,提问作者Adrian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 20:57:07