You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言正则提取含双引号的子名字段问题求助

R语言正则提取含内嵌双引号的子名字段修复方案

原代码

ex02ChildrenInverse <- function(sentence) {
  
  assertString(sentence)
  
  matches <- regmatches(
    
    sentence,
    
    regexec('^(.*?) is the (father|mother) of &quot;(.*?)&quot;', sentence))[[1]]
  
  parent <- matches[[2]]
  
  male <- matches[[3]] == "father"
  
  child <- matches[[4]]
  child <- gsub('&quot;.*&quot;', '', matches[4])
  
  return(list(parent = parent, male = male, child = child))
}

问题说明

当前代码无法正确提取包含内嵌双引号的子名,例如输入:

input <- 'Gudrun is the mother of &quot;Rosamunde (&quot;Rosi&quot;)&quot;.'

当前输出:

$parent
[1] "Gudrun"

$male
[1] FALSE

$child
[1] "Rosamunde ("

期望输出:

$parent
[1] "Gudrun"

$male
[1] FALSE

$child
[1] "Rosamunde (&quot;Rosi&quot;)"

修复方案

问题根源有两点:

  • 原正则使用非贪婪匹配.*?,导致第三个捕获组遇到第一个闭合的&quot;就终止匹配,无法覆盖内嵌双引号的完整子名;
  • 后续的gsub语句逻辑错误,会误删子名中的有效内容。

修改后的代码

ex02ChildrenInverse <- function(sentence) {
  
  assertString(sentence)
  
  # 调整正则:用正向断言匹配到最后一个&quot;前的全部内容
  matches <- regmatches(
    sentence,
    regexec('^(.*?) is the (father|mother) of &quot;(.*)(?=&quot;)', sentence))[[1]]
  
  parent <- matches[[2]]
  male <- matches[[3]] == "father"
  child <- matches[[4]]  # 直接使用捕获到的完整子名
  
  return(list(parent = parent, male = male, child = child))
}

关键修改说明

  1. 将正则中第三个捕获组的.*?改为.*(贪婪匹配),并结合正向断言(?=&quot;)限定匹配边界为最后一个&quot;,确保捕获到从开头到末尾闭合引号前的全部内容;
  2. 移除错误的gsub语句,直接使用捕获组提取的内容作为子名。

如果你的输入中实际使用的是R原生的双引号转义(即\"而非&quot;),只需将正则中的&quot;替换为"即可:

regexec('^(.*?) is the (father|mother) of "(.*)(?=")', sentence)

内容的提问来源于stack exchange,提问作者tong tong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 06:06:13