You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ifelse与dplyr中if_else返回NA时的差异及适配方法

在dplyr中使用if_else处理行内最大值列位置的问题

问题背景

先创建目标数据框:

data.frame(v1 = c(3,2,NA,6),
           v2 = c(NA,1,NA,7),
           v3 = c(1,1,NA,1)) -> df2

需求是按行计算最大值所在的列位置,若整行全为NA则返回NA,而非integer(0)。使用基础R的ifelse可以正常运行:

library(dplyr)
df2 %>% 
  rowwise() %>% 
  mutate(maxCol = ifelse(test = all(is.na(c_across(everything()))),
                         yes = NA, 
                         no = which.max(c_across(everything()))))

但改用dplyr的if_else时会报错:

df2 %>% 
  rowwise() %>% 
  mutate(maxCol = if_else(condition = all(is.na(c_across(everything()))), 
                          true = NA, 
                          false = which.max(c_across(everything()))))

错误提示:false must have size 1, not size 0.

报错原因

dplyr的if_else是严格类型和长度检查的函数:

  • 它会提前计算true和false两个分支的所有结果,而非仅根据条件执行对应分支
  • 当某行全为NA时,which.max(c_across(everything()))会返回integer(0)(长度为0),而true分支的NA长度为1,两者长度不匹配,触发报错
  • 基础R的ifelse是宽松检查逻辑,仅执行满足条件的分支,且会自动调整结果长度,因此不会出现该问题

解决方案

方案1:用case_when替代if_else

case_when对长度的限制更灵活,适配场景需求:

df2 %>% 
  rowwise() %>% 
  mutate(maxCol = case_when(
    all(is.na(c_across(everything()))) ~ NA_integer_,
    TRUE ~ which.max(c_across(everything()))
  ))

注意使用NA_integer_而非NA,保证与which.max返回的整数类型统一。

方案2:在false分支确保返回长度为1的结果

手动处理which.max的空结果,用内部if语句包裹:

df2 %>% 
  rowwise() %>% 
  mutate(maxCol = if_else(
    condition = all(is.na(c_across(everything()))),
    true = NA_integer_,
    false = {
      res <- which.max(c_across(everything()))
      if(length(res) == 0) NA_integer_ else res
    }
  ))

方案3:改用pmap逐行处理

结合purrr的pmap函数,无需rowwise()即可高效逐行计算:

library(purrr)
df2 %>% 
  mutate(maxCol = pmap_int(., ~{
    vals <- c(...)
    if(all(is.na(vals))) NA_integer_ else which.max(vals)
  }))

内容的提问来源于stack exchange,提问作者llewmills

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 20:10:05