如何使用dplyr按组选取每列的首个非NA值
使用dplyr按组选取每列首个非NA值
针对你给出的带NA值的分组tibble,我们可以用dplyr的分组+列操作组合来实现需求:
首先还原你的数据:
library(tibble) library(dplyr) df <- tibble( group = c(rep(1:3, each = 2), 4), x1 = c(1, NA, NA, 4, 5, NA, 7), x2 = c(NA, 100, 30, NA, 3, NA, NA) )
接下来用dplyr处理:
df %>% group_by(group) %>% summarise(across(c(x1, x2), ~first(na.omit(.x))), .groups = "drop")
代码说明:
group_by(group):按group列完成分组across(c(x1, x2), ~first(na.omit(.x))):对指定的x1、x2列执行操作——先用na.omit()剔除该组内的NA值,再用first()取第一个剩余值;如果某组某列全是NA,会自动返回NA.groups = "drop":处理完成后取消分组状态,得到常规结构的tibble
运行后得到目标结果:
# A tibble: 4 × 3 group x1 x2 <dbl> <dbl> <dbl> 1 1 1 100 2 2 4 30 3 3 5 3 4 4 7 NA
如果要对所有非分组列批量操作,无需手动指定列名,可把c(x1, x2)替换为everything():
df %>% group_by(group) %>% summarise(across(everything(), ~first(na.omit(.x))), .groups = "drop")
内容的提问来源于stack exchange,提问作者moremo
相关产品推荐
相关产品推荐

