如何提取字符串第二个下划线后的数字?R语言stringr正则问题
Hey there! Let's figure out why your regex is returning NA and fix it up.
Your original regex (^_){2}(\\d+\\.*\\d*) is trying to match two underscores at the very start of the string (^ anchors to the start), but your input string starts with 1/2/3—that's why it can't find a match and returns NA.
Fix 1: Use a Positive Lookbehind
This approach uses a lookbehind to "skip" everything up to the second underscore, then extracts the digits that follow:
library(stringr) input <- "1/2/3_500k/855kk_1400k/AVBB" # Extract digits after the second underscore result <- str_extract(input, "(?<=^.*?_.*?_)(\\d+\\.*\\d*)") print(result) # Output: [1] "1400"
Breakdown:
(?<=^.*?_.*?_): This positive lookbehind checks that the text before our target digits follows the pattern: start of string → any characters (non-greedily) → underscore → any characters → underscore. This effectively skips everything up to the second underscore.(\\d+\\.*\\d*): Matches integers or decimal numbers (so it works if you ever have something like1400.5instead of1400).
Fix 2: Use str_match to Capture the Target Group
If you prefer a more explicit approach, str_match lets you capture specific parts of the match using groups:
library(stringr) input <- "1/2/3_500k/855kk_1400k/AVBB" # Match up to the second underscore, then capture the following digits match_matrix <- str_match(input, "^.*?_.*?_([0-9.]+)") # Grab the second column (the captured group) result <- match_matrix[, 2] print(result) # Output: [1] "1400"
Breakdown:
^.*?_.*?_: Matches all text from the start of the string up to (and including) the second underscore.([0-9.]+): Captures one or more digits or decimal points—this is our target value.str_matchreturns a matrix where the first column is the full matched text, and subsequent columns are the captured groups. We take the second column to get just the digits.
内容的提问来源于stack exchange,提问作者Alexander
相关产品推荐
相关产品推荐

