R语言正则提取含双引号的子名字段问题求助
R语言正则提取含内嵌双引号的子名字段修复方案
原代码
ex02ChildrenInverse <- function(sentence) { assertString(sentence) matches <- regmatches( sentence, regexec('^(.*?) is the (father|mother) of "(.*?)"', sentence))[[1]] parent <- matches[[2]] male <- matches[[3]] == "father" child <- matches[[4]] child <- gsub('".*"', '', matches[4]) return(list(parent = parent, male = male, child = child)) }
问题说明
当前代码无法正确提取包含内嵌双引号的子名,例如输入:
input <- 'Gudrun is the mother of "Rosamunde ("Rosi")".'
当前输出:
$parent [1] "Gudrun" $male [1] FALSE $child [1] "Rosamunde ("
期望输出:
$parent [1] "Gudrun" $male [1] FALSE $child [1] "Rosamunde ("Rosi")"
修复方案
问题根源有两点:
- 原正则使用非贪婪匹配
.*?,导致第三个捕获组遇到第一个闭合的"就终止匹配,无法覆盖内嵌双引号的完整子名; - 后续的
gsub语句逻辑错误,会误删子名中的有效内容。
修改后的代码
ex02ChildrenInverse <- function(sentence) { assertString(sentence) # 调整正则:用正向断言匹配到最后一个"前的全部内容 matches <- regmatches( sentence, regexec('^(.*?) is the (father|mother) of "(.*)(?=")', sentence))[[1]] parent <- matches[[2]] male <- matches[[3]] == "father" child <- matches[[4]] # 直接使用捕获到的完整子名 return(list(parent = parent, male = male, child = child)) }
关键修改说明
- 将正则中第三个捕获组的
.*?改为.*(贪婪匹配),并结合正向断言(?=")限定匹配边界为最后一个",确保捕获到从开头到末尾闭合引号前的全部内容; - 移除错误的
gsub语句,直接使用捕获组提取的内容作为子名。
如果你的输入中实际使用的是R原生的双引号转义(即\"而非"),只需将正则中的"替换为"即可:
regexec('^(.*?) is the (father|mother) of "(.*)(?=")', sentence)
内容的提问来源于stack exchange,提问作者tong tong
相关产品推荐
相关产品推荐

