You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中提取字符串匹配后的字符?附JSON字段提取实例

Got it, let's tackle this problem step by step. You're on the right track with str_extract from the stringr package—we just need to tweak the regex to target exactly what you want.


First, set up the stringr package

Make sure you have it installed and loaded:

# Install if you haven't already
install.packages("stringr")
# Load the package
library(stringr)

Core Solution: Targeted Regex with str_extract

Your goal is to extract the 14-digit number immediately following {"blur": ". We can use a positive lookbehind assertion to pinpoint this prefix, then match the 14-digit sequence.

Here's the regex you need:

'(?<={"blur": ")\\d{14}'

Let's break down the regex parts:

  • (?<={"blur": "): This is a positive lookbehind—it checks that the text before our target number is exactly {"blur": " (note: in R's single-quoted strings, double quotes don't need escaping).
  • \\d{14}: Matches exactly 14 consecutive digits (\\d represents a digit, {14} enforces the length).

Apply this to your data frame like so:

# Replace `df` with your data frame name, `json_column` with your JSON text column
df$extracted_blur <- str_extract(df$json_column, '(?<={"blur": ")\\d{14}')

Test Example to Verify

Let's create a sample data frame to see how this works:

# Test data
df <- data.frame(json_column = c(
  '{"name": "Bob", "blur": "10010010010010", "age": 30}',
  '{"blur": "98765432109876", "id": 456}',
  '{"status": "active", "note": "No blur value here"}'
))

# Extract the numbers
df$extracted_blur <- str_extract(df$json_column, '(?<={"blur": ")\\d{14}')

# View results
print(df)

You'll get this output:

json_column extracted_blur
1 {"name": "Bob", "blur": "10010010010010", "age": 30} 10010010010010
2                {"blur": "98765432109876", "id": 456} 98765432109876
3           {"status": "active", "note": "No blur value here"}           <NA>

Alternative: Using Capture Groups

If lookbehind assertions feel tricky, you can use a capture group instead. Wrap the 14-digit sequence in parentheses, then use str_match to pull out the captured value:

# Use str_match to get the captured group
matches <- str_match(df$json_column, '{"blur": "(\\d{14})')
df$extracted_blur <- matches[, 2] # The second column holds the captured number

Quick Notes

  • If your JSON uses single quotes instead of double quotes, adjust the regex to use ' instead of ".
  • Observations without the {"blur": " prefix will return NA, which is expected behavior.

内容的提问来源于stack exchange,提问作者Martin M.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:20:25