如何在R语言中查找字符串冒号前的n个单词?优先用正则表达式
Got it, let's break this down using stringr and regular expressions—super straightforward once you map out the pattern! Here's how to pull the n words before that colon in your example string.
First, let's set up our example string and load the package:
library(stringr) input_str <- "Input on income economic activities: Small business, self-emp…"
Scenario 1: Extract the last n words before the colon
This is probably what you're after if you want the closest n words to the colon. We'll use a positive lookahead ((?=:)) to anchor our match right before the colon, then build a pattern to capture n words.
For example, if you want the last 3 words:
n <- 3 # Dynamically build the regex pattern pattern <- str_glue("(\\b\\w+\\b\\s+){{{n-1}}}\\b\\w+\\b(?=:)") # Extract the match str_extract(input_str, pattern) # Output: "income economic activities"
Let's unpack the regex:
\b: Word boundary, ensures we grab full words (not partial chunks)\w+: Matches letters, numbers, or underscores (adjust to[a-zA-Z-]+if your words have hyphens!)(?=:): Positive lookahead—tells regex "only match this if a colon comes right after"str_glue: Lets us dynamically adjust the pattern based onn(no manual regex edits needed!)
Scenario 2: Extract the first n words before the colon
If you want the earliest n words in the text before the colon, we can anchor to the start of the string instead:
n <- 2 pattern <- str_glue("^(\\b\\w+\\b\\s+){{{n-1}}}\\b\\w+\\b") str_extract(input_str, pattern) # Output: "Input on"
Bonus: Two-step approach (more readable!)
If you prefer splitting the task into clear steps, first extract all text before the colon, then pull your n words from that substring:
# Step 1: Grab everything before the colon before_colon <- str_extract(input_str, "^.+?(?=:)") # Step 2: Extract last 2 words from that substring n <- 2 last_n_words <- str_extract(before_colon, str_glue("(\\b\\w+\\b\\s+){{{n-1}}}\\b\\w+\\b$")) # Output: "economic activities" # Or extract first 4 words first_n_words <- str_extract(before_colon, str_glue("^(\\b\\w+\\b\\s+){{{3}}}\\b\\w+\\b")) # Output: "Input on income economic"
Just tweak the regex character class (\w+) if your words include special characters (like hyphens or apostrophes)—for example, use [a-zA-Z'-]+ to match words like "self-employed" or "Mary's".
内容的提问来源于stack exchange,提问作者elliot

