如何在R语言中用正则表达式提取word1与word2前的数字?
Got it, let's work through this problem together! It sounds like you've been trying to pull numbers that come right before word1 and word2, but only managed to grab the first match—totally a common pitfall with regex in R. Here's how you can fix this and capture both numbers:
Key Issue to Fix
The main reason you're only getting the first number is probably because you used a function that stops at the first match (like str_extract or base R's regexpr). To get all matches, you need functions designed to find every occurrence, paired with a regex pattern that targets both word1 and word2.
Solution with stringr (Recommended for Readability)
The stringr package makes regex tasks really straightforward. Use str_extract_all to pull all matching numbers, and a positive lookahead regex to target numbers followed by either word1 or word2.
Here's a complete example:
# Load the stringr package (install it first if you haven't: install.packages("stringr")) library(stringr) # Example test string test_text <- "123 word1 456 word2 789 random_text 101word1" # Regex pattern: matches digits followed by optional whitespace and either word1/word2 pattern <- "\\d+(?=\\s*(?:word1|word2))" # Extract all matches extracted_numbers <- str_extract_all(test_text, pattern)[[1]] # View the result extracted_numbers # Output: [1] "123" "456" "101"
Let's break down the regex pattern:
\\d+: Matches one or more digits (the number we want to capture)(?=...): Positive lookahead—checks that the digits are followed by the content inside, without including that content in the match\\s*: Matches zero or more whitespace characters (handles cases where numbers are right up against the word, like "101word1")(?:word1|word2): A non-capturing group that matches eitherword1orword2(the?:tells regex not to save this group as a separate match)
Solution with Base R
If you prefer not to use external packages, you can use gregexpr (which finds all match positions) and regmatches (which extracts the matches). Note that you need to enable Perl-compatible regex for lookaheads to work:
test_text <- "123 word1 456 word2 789 random_text 101word1" # Use gregexpr with perl=TRUE to support lookaheads match_positions <- gregexpr("\\d+(?=\\s*(?:word1|word2))", test_text, perl = TRUE) extracted_numbers <- regmatches(test_text, match_positions)[[1]] extracted_numbers # Output: [1] "123" "456" "101"
Why This Works
By using functions that capture all matches (str_extract_all or gregexpr) and a pattern that explicitly targets both word1 and word2, you'll get every number that precedes either of your target words—no more missing the second match!
内容的提问来源于stack exchange,提问作者Stefania Axo

