基于正则实现grepl/gsub忽略指定分隔符内子串的方案问询
Great question! The placeholder-based approach you're currently using can get shaky fast—especially with messy real-world text. Instead, we can use Perl-compatible regex tricks to directly target matches only outside your specified delimiter pairs, no placeholders required. Here's a robust implementation:
Helper Function: Parse Delimiters
First, we'll add a helper to safely escape regex-special characters in your delimiters (like {, }, ", etc.):
parse_delimiters <- function(escaped) { lapply(strsplit(escaped, ""), function(pair) { list( open = gsub("([][{}().+*^$|\\?])", "\\\\\\1", pair[1]), close = gsub("([][{}().+*^$|\\?])", "\\\\\\1", pair[2]) ) }) }
grepl2() Implementation
This version uses (*SKIP)(*FAIL) to ignore text inside your delimiter pairs, then checks for your target pattern in the remaining text:
grepl2 <- function(pattern, x, ignore.case = FALSE, perl = TRUE, fixed = FALSE, useBytes = FALSE, escaped = c('""', "''")) { if (!perl) stop("This implementation requires `perl=TRUE` for regex skip/fail support") if (fixed) stop("`fixed=TRUE` isn't compatible with this regex approach—use your original placeholder version if you need fixed matching") delimiters <- parse_delimiters(escaped) # Build patterns to skip text inside each delimiter pair skip_patterns <- vapply(delimiters, function(d) { sprintf("(%s.*?%s)(*SKIP)(*FAIL)", d$open, d$close) }, character(1)) # Combine skip rules with the target pattern full_pattern <- paste(c(skip_patterns, pattern), collapse = "|") grepl(full_pattern, x, ignore.case = ignore.case, perl = TRUE, useBytes = useBytes) }
gsub2() Implementation
The replacement function works the same way—we skip delimited text, then only replace matches in the remaining content:
gsub2 <- function(pattern, replacement, x, ignore.case = FALSE, perl = TRUE, fixed = FALSE, useBytes = FALSE, escaped = c('""', "''")) { if (!perl) stop("This implementation requires `perl=TRUE` for regex skip/fail support") if (fixed) stop("`fixed=TRUE` isn't compatible with this regex approach—use your original placeholder version if you need fixed matching") delimiters <- parse_delimiters(escaped) # Build patterns to skip text inside each delimiter pair skip_patterns <- vapply(delimiters, function(d) { sprintf("(%s.*?%s)(*SKIP)(*FAIL)", d$open, d$close) }, character(1)) # Combine skip rules with the target pattern full_pattern <- paste(c(skip_patterns, pattern), collapse = "|") gsub(full_pattern, replacement, x, ignore.case = ignore.case, perl = TRUE, useBytes = useBytes) }
How This Works
(*SKIP)(*FAIL)Regex Trick: When the regex matches text between your open/close delimiters,(*SKIP)tells the engine not to backtrack into that text, and(*FAIL)forces that match to be discarded. Only text outside delimiter pairs is considered for your target pattern.- Safe Delimiter Handling: The
parse_delimitersfunction ensures special regex characters in your delimiters (like{or$) are treated as literal characters, avoiding unexpected behavior. - Perl Requirement: This relies on PCRE features, so
perl=TRUEis mandatory. We've added checks to enforce this and warn about incompatible options likefixed=TRUE.
Test Your Sample Cases
Let's verify that all your example scenarios work as expected:
# grepl2 tests grepl2("banana", "'banana' banana \"banana\"", escaped = c('""', "''")) #> [1] TRUE grepl2("banana", "'banana' apple \"banana\"", escaped = c('""', "''")) #> [1] FALSE grepl2("banana", "{banana} banana {banana}", escaped = "{}") #> [1] TRUE grepl2("banana", "{banana} apple {banana}", escaped = "{}") #> [1] FALSE # gsub2 tests gsub2("banana", "potatoe", "'banana' banana \"banana\"") #> [1] "'banana' potatoe \"banana\"" gsub2("banana", "potatoe", "'banana' apple \"banana\"") #> [1] "'banana' apple \"banana\"" gsub2("banana", "potatoe", "{banana} banana {banana}", escaped = "{}") #> [1] "{banana} potatoe {banana}" gsub2("banana", "potatoe", "{banana} apple {banana}", escaped = "{}") #> [1] "{banana} apple {banana}"
All tests pass perfectly!
Key Notes
- Balanced Delimiters: This assumes your delimiter pairs are properly balanced and non-overlapping (e.g., no
'{"}'cases). For nested or unbalanced delimiters, you'd need a more complex parser (but this covers most common use cases). - Case Sensitivity: The
ignore.caseparameter works as expected, applying to both your target pattern and delimiter matches (though delimiters are usually case-sensitive in practice). - Byte-wise Matching: The
useBytesparameter is passed through togrepl/gsubfor handling byte-wise matching if needed.
内容的提问来源于stack exchange,提问作者moodymudskipper

