You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用dplyr按组选取每列的首个非NA值

使用dplyr按组选取每列首个非NA值

针对你给出的带NA值的分组tibble,我们可以用dplyr的分组+列操作组合来实现需求:

首先还原你的数据:

library(tibble)
library(dplyr)

df <- tibble(
  group = c(rep(1:3, each = 2), 4),
  x1 = c(1, NA, NA, 4, 5, NA, 7),
  x2 = c(NA, 100, 30, NA, 3, NA, NA)
)

接下来用dplyr处理:

df %>%
  group_by(group) %>%
  summarise(across(c(x1, x2), ~first(na.omit(.x))), .groups = "drop")

代码说明:

  • group_by(group):按group列完成分组
  • across(c(x1, x2), ~first(na.omit(.x))):对指定的x1、x2列执行操作——先用na.omit()剔除该组内的NA值,再用first()取第一个剩余值;如果某组某列全是NA,会自动返回NA
  • .groups = "drop":处理完成后取消分组状态,得到常规结构的tibble

运行后得到目标结果:

# A tibble: 4 × 3
  group    x1    x2
  <dbl> <dbl> <dbl>
1     1     1   100
2     2     4    30
3     3     5     3
4     4     7    NA

如果要对所有非分组列批量操作,无需手动指定列名,可把c(x1, x2)替换为everything():

df %>%
  group_by(group) %>%
  summarise(across(everything(), ~first(na.omit(.x))), .groups = "drop")

内容的提问来源于stack exchange,提问作者moremo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 09:02:10