如何在R语言中计算指定列的行最小值(排除0和NA)
问题需求
基于R语言的dplyr包,为数据框新增一列nieuw,规则如下:
- 取指定3列(
test1/test2/test3)的行最小值,排除0和NA - 若行内仅包含0和NA,返回0
- 若行内仅包含NA,返回NA
现有解决方案逻辑复杂且运行时产生警告,需要更简洁的实现方式。
原数据与现有代码运行结果:
library(dplyr) df <- data.frame( id = c(1, 2, 3, 4), test1 = c( NA, NA, 2 , 3), test2 = c( NA, 0, 1 , 1), test3 = c(NA, NA, 0 , 2) ) # 原代码(有警告,逻辑繁琐) df2 <- df %>% mutate(nieuw = apply(across(test1:test3), 1, function(x) min(x[x>0]))) %>% rowwise() %>% mutate(nieuw = if_else(is.na(nieuw), max(across(test1:test3), na.rm = TRUE), nieuw)) %>% mutate(nieuw = ifelse(is.infinite(nieuw), NA, nieuw))
运行结果:
> df2 # A tibble: 4 x 5 # Rowwise: id test1 test2 test3 nieuw <dbl> <dbl> <dbl> <dbl> <dbl> 1 1 NA NA NA NA 2 2 NA 0 NA 0 3 3 2 1 0 1 4 4 3 1 2 1 Warning message: Problem while computing `nieuw = if_else(...)`. i no non-missing arguments to max; returning -Inf i The warning occurred in row 1.
简洁解决方案
方法1:封装自定义函数(可读性最高)
先定义一个处理单行逻辑的函数,再结合rowwise()逐行应用:
# 定义处理单行的函数 min_excl0_na <- function(x) { # 筛选出大于0的有效数值 valid_vals <- x[x > 0] # 分支判断三种情况 if (length(valid_vals) > 0) { min(valid_vals, na.rm = TRUE) } else if (all(is.na(x))) { NA_real_ } else { 0 } } # 应用到数据框 df_clean <- df %>% rowwise() %>% mutate(nieuw = min_excl0_na(c(test1, test2, test3))) %>% ungroup() # 取消行分组,避免后续操作受影响
方法2:inline逻辑(无需额外函数)
直接在mutate中用case_when完成所有判断,更紧凑:
df_clean <- df %>% rowwise() %>% mutate( # 临时存储当前行的有效数值(>0) valid_vals = list(c(test1, test2, test3)[c(test1, test2, test3) > 0]), nieuw = case_when( length(valid_vals) > 0 ~ min(valid_vals), all(is.na(c(test1, test2, test3))) ~ NA_real_, TRUE ~ 0 ) ) %>% select(-valid_vals) %>% # 移除临时列 ungroup()
效果验证
运行后得到无警告的正确结果:
> df_clean # A tibble: 4 x 5 id test1 test2 test3 nieuw <dbl> <dbl> <dbl> <dbl> <dbl> 1 1 NA NA NA NA 2 2 NA 0 NA 0 3 3 2 1 0 1 4 4 3 1 2 1
优势对比
- 完全消除原代码中因
max(..., na.rm=TRUE)处理全NA行产生的警告 - 逻辑集中在一处,可读性和可维护性更强
- 避免多次
mutate嵌套,代码更简洁
内容的提问来源于stack exchange,提问作者Nina van Bruggen
相关产品推荐
相关产品推荐

