在mutate中使用case_when处理缺失变量的技术问题
解决dplyr::case_when中变量不存在时的报错问题
问题核心在于dplyr::case_when会提前评估所有分支的表达式,即使左侧条件已判断变量不存在,依然会尝试解析右侧的变量引用,导致变量缺失时报错。以下是两种可行的解决方案:
方案一:先统一变量存在性(推荐)
先为可能缺失的变量创建临时副本,不存在则设为NA,处理完成后再移除临时添加的变量,逻辑清晰且兼容两种场景:
df_mut <- function(df) { # 记录原始数据中是否存在目标变量 has_txt <- "info_1_txt" %in% colnames(df) has_dtl <- "info_1_dtl" %in% colnames(df) res <- df |> # 确保变量存在,不存在则设为NA字符型 dplyr::mutate( info_1_txt = if (has_txt) info_1_txt else NA_character_, info_1_dtl = if (has_dtl) info_1_dtl else NA_character_ ) |> # 生成info_1字段 dplyr::mutate( info_1 = dplyr::case_when( !is.na(info_1_txt) & !is.na(info_1_dtl) ~ paste0(info_1_txt, "|", info_1_dtl), !is.na(info_1_dtl) ~ info_1_dtl, !is.na(info_1_txt) ~ info_1_txt, .default = NA_character_ ) ) |> # 移除临时添加的变量(仅当原始数据中不存在时) dplyr::select(-dplyr::all_of(c("info_1_txt"[!has_txt], "info_1_dtl"[!has_dtl]))) res }
方案二:在case_when中动态提取变量
通过dplyr::pull结合存在性判断,仅当变量存在时才尝试提取,避免直接引用不存在的变量:
df_mut <- function(df) { res <- df |> dplyr::mutate( info_1 = dplyr::case_when( "info_1_txt" %in% colnames(df) & "info_1_dtl" %in% colnames(df) & !is.na(dplyr::pull(df, info_1_txt)) & !is.na(dplyr::pull(df, info_1_dtl)) ~ paste0(dplyr::pull(df, info_1_txt), "|", dplyr::pull(df, info_1_dtl)), "info_1_dtl" %in% colnames(df) & !is.na(dplyr::pull(df, info_1_dtl)) ~ dplyr::pull(df, info_1_dtl), "info_1_txt" %in% colnames(df) ~ dplyr::pull(df, info_1_txt), .default = NA_character_ ) ) res }
验证效果
测试变量存在的场景
df_1 <- tibble::tibble( last_name = c("Doe", "Doe"), first_name = c("John", "Jane"), info_1_txt = c("economics", "culture"), info_1_dtl = c("glare", "deficiency") ) df_1_res <- df_mut(df_1) |> dplyr::glimpse()
输出:
Rows: 2 Columns: 5 $ last_name <chr> "Doe", "Doe" $ first_name <chr> "John", "Jane" $ info_1_txt <chr> "economics", "culture" $ info_1_dtl <chr> "glare", "deficiency" $ info_1 <chr> "economics|glare", "culture|deficiency"
测试变量缺失的场景
df_2 <- tibble::tibble( last_name = c("Smith", "Smith"), first_name = c("Tom", "Dick") ) df_2_res <- df_mut(df_2) |> dplyr::glimpse()
输出符合预期:
Rows: 2 Columns: 3 $ last_name <chr> "Smith", "Smith" $ first_name <chr> "Tom", "Dick" $ info_1 <chr> NA, NA
内容的提问来源于stack exchange,提问作者pure_func
相关产品推荐
相关产品推荐

