如何在R语言中提取字符串匹配后的字符?附JSON字段提取实例
Got it, let's tackle this problem step by step. You're on the right track with str_extract from the stringr package—we just need to tweak the regex to target exactly what you want.
First, set up the stringr package
Make sure you have it installed and loaded:
# Install if you haven't already install.packages("stringr") # Load the package library(stringr)
Core Solution: Targeted Regex with str_extract
Your goal is to extract the 14-digit number immediately following {"blur": ". We can use a positive lookbehind assertion to pinpoint this prefix, then match the 14-digit sequence.
Here's the regex you need:
'(?<={"blur": ")\\d{14}'
Let's break down the regex parts:
(?<={"blur": "): This is a positive lookbehind—it checks that the text before our target number is exactly{"blur": "(note: in R's single-quoted strings, double quotes don't need escaping).\\d{14}: Matches exactly 14 consecutive digits (\\drepresents a digit,{14}enforces the length).
Apply this to your data frame like so:
# Replace `df` with your data frame name, `json_column` with your JSON text column df$extracted_blur <- str_extract(df$json_column, '(?<={"blur": ")\\d{14}')
Test Example to Verify
Let's create a sample data frame to see how this works:
# Test data df <- data.frame(json_column = c( '{"name": "Bob", "blur": "10010010010010", "age": 30}', '{"blur": "98765432109876", "id": 456}', '{"status": "active", "note": "No blur value here"}' )) # Extract the numbers df$extracted_blur <- str_extract(df$json_column, '(?<={"blur": ")\\d{14}') # View results print(df)
You'll get this output:
json_column extracted_blur 1 {"name": "Bob", "blur": "10010010010010", "age": 30} 10010010010010 2 {"blur": "98765432109876", "id": 456} 98765432109876 3 {"status": "active", "note": "No blur value here"} <NA>
Alternative: Using Capture Groups
If lookbehind assertions feel tricky, you can use a capture group instead. Wrap the 14-digit sequence in parentheses, then use str_match to pull out the captured value:
# Use str_match to get the captured group matches <- str_match(df$json_column, '{"blur": "(\\d{14})') df$extracted_blur <- matches[, 2] # The second column holds the captured number
Quick Notes
- If your JSON uses single quotes instead of double quotes, adjust the regex to use
'instead of". - Observations without the
{"blur": "prefix will returnNA, which is expected behavior.
内容的提问来源于stack exchange,提问作者Martin M.

