You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在mutate中使用case_when处理缺失变量的技术问题

解决dplyr::case_when中变量不存在时的报错问题

问题核心在于dplyr::case_when会提前评估所有分支的表达式,即使左侧条件已判断变量不存在,依然会尝试解析右侧的变量引用,导致变量缺失时报错。以下是两种可行的解决方案:

方案一:先统一变量存在性(推荐)

先为可能缺失的变量创建临时副本,不存在则设为NA,处理完成后再移除临时添加的变量,逻辑清晰且兼容两种场景:

df_mut <- function(df) {
  # 记录原始数据中是否存在目标变量
  has_txt <- "info_1_txt" %in% colnames(df)
  has_dtl <- "info_1_dtl" %in% colnames(df)
  
  res <- df |> 
    # 确保变量存在,不存在则设为NA字符型
    dplyr::mutate(
      info_1_txt = if (has_txt) info_1_txt else NA_character_,
      info_1_dtl = if (has_dtl) info_1_dtl else NA_character_
    ) |>
    # 生成info_1字段
    dplyr::mutate(
      info_1 = dplyr::case_when(
        !is.na(info_1_txt) & !is.na(info_1_dtl) ~ paste0(info_1_txt, "|", info_1_dtl),
        !is.na(info_1_dtl) ~ info_1_dtl,
        !is.na(info_1_txt) ~ info_1_txt,
        .default = NA_character_
      )
    ) |>
    # 移除临时添加的变量(仅当原始数据中不存在时)
    dplyr::select(-dplyr::all_of(c("info_1_txt"[!has_txt], "info_1_dtl"[!has_dtl])))
  
  res
}

方案二:在case_when中动态提取变量

通过dplyr::pull结合存在性判断,仅当变量存在时才尝试提取,避免直接引用不存在的变量:

df_mut <- function(df) {
  res <- df |> 
    dplyr::mutate(
      info_1 = dplyr::case_when(
        "info_1_txt" %in% colnames(df) & "info_1_dtl" %in% colnames(df) &
          !is.na(dplyr::pull(df, info_1_txt)) & !is.na(dplyr::pull(df, info_1_dtl)) ~
          paste0(dplyr::pull(df, info_1_txt), "|", dplyr::pull(df, info_1_dtl)),
        "info_1_dtl" %in% colnames(df) & !is.na(dplyr::pull(df, info_1_dtl)) ~
          dplyr::pull(df, info_1_dtl),
        "info_1_txt" %in% colnames(df) ~ dplyr::pull(df, info_1_txt),
        .default = NA_character_
      )
    )
  res
}

验证效果

测试变量存在的场景

df_1 <- 
  tibble::tibble(
    last_name = c("Doe", "Doe"),
    first_name = c("John", "Jane"),
    info_1_txt = c("economics", "culture"),
    info_1_dtl = c("glare", "deficiency")
  )

df_1_res <- df_mut(df_1) |> dplyr::glimpse()

输出:

Rows: 2
Columns: 5
$ last_name  <chr> "Doe", "Doe"
$ first_name <chr> "John", "Jane"
$ info_1_txt <chr> "economics", "culture"
$ info_1_dtl <chr> "glare", "deficiency"
$ info_1     <chr> "economics|glare", "culture|deficiency"

测试变量缺失的场景

df_2 <- 
  tibble::tibble(
    last_name = c("Smith", "Smith"),
    first_name = c("Tom", "Dick")
  )

df_2_res <- df_mut(df_2) |> dplyr::glimpse()

输出符合预期:

Rows: 2
Columns: 3
$ last_name  <chr> "Smith", "Smith"
$ first_name <chr> "Tom", "Dick"
$ info_1     <chr> NA, NA

内容的提问来源于stack exchange,提问作者pure_func

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 04:52:15