如何利用stringr包提取字符串中所有以1.开头的匹配模式?
Hey there! I get it—regex can feel overwhelming at first, but let's walk through exactly how to extract all the numbers and "1."-starting patterns you need using R's stringr package. The key here is swapping out str_extract (which only grabs the first match) for str_extract_all, which pulls every matching instance.
Step 1: Load the stringr Package
First, make sure you have the package loaded (install it with install.packages("stringr") if you haven't already):
library(stringr)
Step 2: Extract All Numbers from a String
Whether you need integers or decimal numbers, here are two common scenarios:
Extract All Integers
Use the regex pattern \\d+ to match one or more consecutive digits:
# Example text sample_text <- "I have 5 apples, 12 oranges, and 3 bananas. Also, 1.2kg of grapes." # Extract all integers all_integers <- str_extract_all(sample_text, "\\d+")[[1]] all_integers # Output: ["5", "12", "3", "1", "2"]
\\d: Regex shorthand for any digit (0-9)+: Matches one or more of the preceding element (so it grabs full numbers, not single digits)[[1]]:str_extract_allreturns a list—since we're working with one text string, we pull the first (and only) element to get a character vector.
Extract All Numbers (Including Decimals)
If you need to capture decimal values too, use \\d+\\.?\\d*:
all_numbers <- str_extract_all(sample_text, "\\d+\\.?\\d*")[[1]] all_numbers # Output: ["5", "12", "3", "1.2"]
\\.?: Matches an optional decimal point (the?means "zero or one of these")\\d*: Matches zero or more digits after the decimal point (so it works for integers and decimals alike)
Step 3: Extract All Patterns Starting with "1."
Let's say you want to grab every section in your text that starts with 1. (like numbered list items). The regex here needs to be specific enough to stop at the next numbered item or the end of the string.
Example Code
list_text <- "1. Buy groceries 2. Walk the dog 1. Finish R project 3. Call mom" # Regex to match "1." followed by text until the next numbered item or end pattern_1_start <- "1\\..*?(?=\\s+\\d\\.|$)" matches_1_start <- str_extract_all(list_text, pattern_1_start)[[1]] matches_1_start # Output: ["1. Buy groceries", "1. Finish R project"]
Breakdown of the Regex:
1\\.: Matches the literal1.—we escape the dot with\\because in regex, a plain.matches any character..*?: The.*matches any character (except newlines) zero or more times, and the?makes it "non-greedy"—this means it stops as soon as it hits the next part of the pattern, instead of grabbing everything until the end of the string.(?=\\s+\\d\\.|$): This is a "positive lookahead"—it checks that the text we're matching is followed by either:\\s+\\d\\.: One or more spaces, a digit, and a dot (the start of another numbered item like2.), or$: The end of the string.
Quick Tip: Handling Multiple Text Strings
If you're working with a vector of multiple text strings, str_extract_all will return a list where each element corresponds to the matches from one string. For example:
text_vector <- c("1. A 2. B", "1. C 3. D") str_extract_all(text_vector, "1\\..*?(?=\\s+\\d\\.|$)") # Output: list(["1. A"], ["1. C"])
内容的提问来源于stack exchange,提问作者Alexander

