在dplyr管道中使用which.min函数遇问题,求更简洁优雅解法
在dplyr管道中优雅处理
which.min缺失值问题 问题说明
使用dplyr管道调用which.min提取每行最小值对应的列索引时,全NA行会触发报错(which.min返回长度为0的向量)。现有条件判断的解决方案较为繁琐,以下提供更简洁的实现方式。
可复现示例
library(dplyr) data <- data.frame(s1=c(10,NA,5,NA,NA), s2=c(8,NA,NA,4,20), s3=c(NA,NA,2,NA,10)) data #> s1 s2 s3 #> 1 10 8 NA #> 2 NA NA NA #> 3 5 NA 2 #> 4 NA 4 NA #> 5 NA 20 10
提取最小值(正常工作)
min(x, na.rm=TRUE)可正常提取每行最小值,仅全NA行返回Inf并给出警告:
data %>% rowwise() %>% mutate(Min_s = min(c(s1,s2,s3), na.rm=TRUE)) #> Warning: There was 1 warning in `mutate()`. #> ℹ In argument: `Min_s = min(c(s1, s2, s3), na.rm = TRUE)`. #> ℹ In row 2. #> Caused by warning in `min()`: #> ! no non-missing arguments to min; returning Inf #> # A tibble: 5 × 4 #> # Rowwise: #> s1 s2 s3 Min_s #> <dbl> <dbl> <dbl> <dbl> #> 1 10 8 NA 8 #> 2 NA NA NA Inf #> 3 5 NA 2 2 #> 4 NA 4 NA 4 #> 5 NA 20 10 10
直接调用which.min的报错问题
全NA行中which.min返回长度为0的向量,不符合dplyr对列长度的要求,触发报错:
data %>% rowwise() %>% mutate(which_s = which.min(c(s1,s2,s3))) #> Error in `mutate()`: #> ℹ In argument: `which_s = which.min(c(s1, s2, s3))`. #> ℹ In row 2. #> Caused by error: #> ! `which_s` must be size 1, not 0. #> ℹ Did you mean: `which_s = list(which.min(c(s1, s2, s3)))` ?
现有繁琐解决方案
通过手动条件判断规避全NA行报错:
data %>% rowwise() %>% mutate(which_s = if(!is.na(s1)|!is.na(s2)|!is.na(s3)) {which.min(c(s1,s2,s3))} else NA ) #> # A tibble: 5 × 4 #> # Rowwise: #> s1 s2 s3 which_s #> <dbl> <dbl> <dbl> <int> #> 1 10 8 NA 2 #> 2 NA NA NA NA #> 3 5 NA 2 3 #> 4 NA 4 NA 2 #> 5 NA 20 10 3
更简洁的优雅实现
方法1:用if_all简化条件判断
用if_all批量检查是否全为NA,替代手动多列判断:
data %>% rowwise() %>% mutate(which_s = if(!if_all(c(s1, s2, s3), is.na)) which.min(c(s1, s2, s3)) else NA)
方法2:用purrr::map_int替代rowwise
避免rowwise的性能开销,通过map_int逐行处理:
library(purrr) data %>% mutate(which_s = pmap_int(across(s1:s3), ~{ vals <- c(...) all(is.na(vals)) %>% ifelse(NA_integer_, which.min(vals)) }))
方法3:用replace处理全NA情况
直接在which.min结果上替换全NA行的值:
data %>% rowwise() %>% mutate(which_s = replace(which.min(c(s1,s2,s3)), all(is.na(c(s1,s2,s3))), NA))
内容的提问来源于stack exchange,提问作者Wael
相关产品推荐
相关产品推荐

