R语言正则提取文本报错:无法提取括号内数字问题咨询
Let's break down the issues with your original code and fix them step by step.
1. Key Errors in Your Original Code
a. Leading .+ Causes Pattern Mismatch
Your regex starts with .+, which means "one or more of any character" before the first (. But your input string starts with (, so there's nothing to match here! This makes the entire regex fail to recognize your string, so gsub just returns the original text instead of modifying it.
b. Unescaped Backreferences
In R, regex backreferences like \1 need to be written as \\1 in string literals. Using \1 tells R to interpret it as a control character (not a regex reference), so even if the pattern matched, the replacement wouldn't work correctly.
2. Fixed Code
Let's adjust the regex to match your input structure and fix the backreferences. Here's the corrected version:
For a Single String
txt = "(2) 1G–1G (0)" gsub('.*\\(([0-9]+)\\) 1G–1G \\(([0-9]+)\\).*', '\\1 - \\2', txt) # Returns: "2 - 0"
For Your Data Frame
DF <- data.frame(txt = c('(2) 1G–1G (0)','(1) 1G–1G (4)','(2) 1G–1G (0)')) DF$result <- gsub('.*\\(([0-9]+)\\) 1G–1G \\(([0-9]+)\\).*', '\\1 - \\2', DF$txt) # Resulting DF: # txt result # 1 (2) 1G–1G (0) 2 - 0 # 2 (1) 1G–1G (4) 1 - 4 # 3 (2) 1G–1G (0) 2 - 0
3. What the Fixed Regex Does
Let's break down the pattern .*\\(([0-9]+)\\) 1G–1G \\(([0-9]+)\\).*:
.*: Matches zero or more characters (handles cases where there might be text before/after your target pattern, or none like in your input).\\(([0-9]+)\\): Captures the first number in parentheses (the([0-9]+)is group 1, and\\(/\\)escape the literal parentheses).1G–1G: Matches the exact literal string between the two parenthetical numbers.\\(([0-9]+)\\): Captures the second number in parentheses (group 2)..*: Matches any remaining characters after the second parenthesis.
The replacement string \\1 - \\2 uses the two captured groups, separated by " - " to get your desired format.
Alternative (More Readable) Approach with stringr
If you're open to using the stringr package, you can use str_match to directly extract the groups, then paste them together:
library(stringr) # Extract captured groups matches <- str_match(DF$txt, '\\(([0-9]+)\\) 1G–1G \\(([0-9]+)\\)') DF$result <- paste(matches[,2], matches[,3], sep = " - ")
This achieves the same result but might be easier to debug if you need to adjust the pattern later.
内容的提问来源于stack exchange,提问作者LeMarque

