如何在R中编写正则匹配指定门牌号到邮政编码间的文本
Got it, let's sort out that regex issue you're facing in R! The problem with your original pattern is indeed the greedy matching causing it to run wild—here's how to fix it and get the exact extraction you need.
The Fixed Regex
First, here's the adjusted pattern that avoids infinite matching and correctly captures the text between your target door number and postal code:
pattern <- "\\b97\\b(?s)(.*?)(3000 AB)"
Or if you prefer a version that doesn't rely on the single-line modifier (for broader compatibility):
pattern <- "\\b97\\b([\\s\\S]*?)(3000 AB)"
What Each Part Does
Let's break this down so you understand why it works:
\\b97\\b: The word boundaries (\\b) ensure we match only the standalone number 97—no partial matches like 197 or 971 will trigger this.(?s): This enables "single-line mode" in R's PCRE regex engine, which makes the.character match all characters including newlines (critical for your real-world scenario with line breaks and random spacing)..*?: The*?is a non-greedy quantifier—unlike the greedy*in your original pattern, it stops matching as soon as it hits the first occurrence of "3000 AB", preventing endless backtracking and infinite loops.(3000 AB): Our target end marker, which signals where to stop extracting text.
Example Usage in R
Here's how to implement this with both stringr (tidyverse) and base R, using text with line breaks to simulate your real-world scenario:
Using stringr (tidyverse)
library(stringr) sample_text <- "Company X Fakestreet 97, This is an invoice. Please pay :) 3000 AB Fakecity" # Extract just the text between 97 and 3000 AB using capture groups middle_text <- str_match(sample_text, pattern)[, 2] print(middle_text) # Output: ",\nThis is an invoice.\nPlease pay :)\n"
Using Base R
sample_text <- "Company X Fakestreet 97, This is an invoice. Please pay :) 3000 AB Fakecity" # Find matches using regexec matches <- regmatches(sample_text, regexec(pattern, sample_text)) # Extract the middle capture group if a match is found if (length(matches[[1]]) > 1) { cat("Extracted text:\n", matches[[1]][2], "\n") }
Why Your Original Regex Failed
Your original pattern \\b(97){1}\\b((.| | | |))*(3000 AB) had two key issues:
- Greedy matching: The
*quantifier tries to match as much text as possible, leading to excessive backtracking—especially with long text or lots of line breaks, this can cause the regex engine to get stuck in loops. - Redundant character matching: The
((.| | | |))*is a messy way to match line breaks; using(?s).*?or[\\s\\S]*?is cleaner and more reliable.
That should solve your problem! Let me know if you need further tweaks for edge cases.
内容的提问来源于stack exchange,提问作者rgms

