使用tidyverse修正NA替换及不匹配行移除的代码问题
问题描述
示例数据集
# Load the tidyverse package library(tidyverse) # Create the dataset id <- 1:6 model <- c("0RB3211", NA, "0RB4191", NA, "0RB4033", NA) UPC <- c("805289119081", "DK_0RB3447CP_RBCP 50", "8053672006360", "Green_Classic_G-15_Polar_1.67_PREM_SV", "805289044604", "DK_0RB2132CP_RBCP 55") df <- tibble(id, model, UPC)
需求说明
当model列存在缺失值(NA)时:
- 若对应
UPC以DK开头,提取第一个下划线后的7位字符数字组合(如第二行提取"0RB3447"、最后一行提取"0RB2132")填充model; - 若
UPC不以DK开头则删除整行(如第四行需删除)。
错误代码
# Manipulate the dataset df_cleaned <- df %>% rowwise() %>% mutate(model = ifelse(is.na(model) & str_detect(UPC, "^DK"), str_extract(UPC, "\\d{2}RB\\d{4}"), model)) %>% ungroup() %>% filter(!(is.na(model) & str_detect(UPC, "[^0-9]"))) # Display the cleaned dataset print(df_cleaned)
请问如何修改上述代码以得到正确结果?
修正方案
原代码问题分析
- 正则表达式错误:原代码用
\\d{2}RB\\d{4}匹配目标字符串,但实际目标是0RB3447(1个数字+RB+4个数字,共7位),该正则无法正确捕获内容。 - 过滤逻辑错误:原过滤条件
!(is.na(model) & str_detect(UPC, "[^0-9]"))逻辑混乱,无法准确筛选需要保留的行。
修正后代码
df_cleaned <- df %>% # 填充NA的model:仅当model为NA且UPC以DK开头时,提取DK_后7位 mutate(model = case_when( is.na(model) & str_detect(UPC, "^DK") ~ str_extract(UPC, "(?<=DK_).{7}"), TRUE ~ model )) %>% # 删除model仍为NA的行(即原NA且UPC不以DK开头的行) filter(!is.na(model)) print(df_cleaned)
代码说明
case_when替代ifelse:逻辑分支更清晰,明确处理特定条件下的填充操作。- 正则表达式优化:
(?<=DK_).{7}是正向预查,精准匹配DK_之后的任意7个字符;如果需要限定只能是字母数字组合,可改为(?<=DK_)[A-Z0-9]{7}。 - 过滤逻辑简化:直接过滤掉
model仍为NA的行,完全符合需求中“删除UPC不以DK开头且model为NA的行”的要求。
内容的提问来源于stack exchange,提问作者Fox_Summer
相关产品推荐
相关产品推荐

