R语言双条件填充new列:字段提取正则匹配问题
new Column Let's tackle this problem step by step. You've already nailed the first part (using values from column b when available), so we just need to fix the regex extraction for when b is empty.
The Core Issue with Your Regex
Your previous attempts didn't work because:
- The first
gsubpattern targets underscores and dot-separated endings, which don't align with your actual data structure. - The second
sublooks forTASK=(with an equals sign), but your data usesTASK(a space) instead—so it never finds a valid match.
Correct Approach
We'll use ifelse() to handle the two cases, and a precise regex to pull the content between TASK and the next space when b is empty.
Base R Solution
Here's the complete code that delivers your desired output:
# Your original data a <- c("11-010 Bla", "TASK 21 MMM", "TASK 03-11-11 Hah") b <- c("11-010","","") df <- data.frame(a, b, new = "") # Populate the new column df$new <- ifelse( df$b != "", # Check if column b has a value df$b, # Use b's value if yes sub(".*TASK (\\S+).*", "\\1", df$a) # Extract from a if no )
Let's break down the regex .*TASK (\\S+).*:
.*: Matches any characters leading up toTASK(\\S+): Captures one or more non-space characters (this is the value we want).*: Matches any remaining characters after the captured value\\1: Replaces the entire string with the captured group
Alternative with stringr (More Readable)
If you're using the stringr package, you can use a positive lookbehind for even cleaner code:
library(stringr) df$new <- ifelse( df$b != "", df$b, str_extract(df$a, "(?<=TASK )\\S+") )
The (?<=TASK ) part is a positive lookbehind—it tells str_extract to find the \\S+ (non-space characters) that come right after TASK .
Final Output
After running either code, your df$new will be exactly what you wanted:"11-010", "21", "03-11-11"
内容的提问来源于stack exchange,提问作者Kalenji

