You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr::across按组统计多列合并后的唯一值数量?

按组统计多列合并后的唯一值数量(dplyr实现)

你的原代码across(DX1:DX4, n_distinct)是对每一列单独计算唯一值,无法实现多列合并后统计唯一值的需求。以下是两种可行的解决方案:

方法一:使用c_across(推荐)

c_across可以将指定列的所有元素合并为一个向量,配合n_distinct直接按组统计唯一值数量:

library(dplyr)

df <- read.table(header = TRUE, text = "
id DX1 DX2 DX3 DX4
1 A B A A
1 A A A C
1 D A A A
1 A A A F
1 A A A A
2 A A A A
2 A C A A
2 A A A D
2 A E D B
", stringsAsFactors = FALSE)

# 按组统计DX1-DX4合并后的唯一值数量
result <- df %>%
  group_by(id) %>%
  summarize(unique_dx_count = n_distinct(c_across(DX1:DX4), na.rm = TRUE))

print(result)

运行结果:

# A tibble: 2 × 2
     id unique_dx_count
  <int>           <int>
1     1               5
2     2               5

方法二:先转长格式再统计

通过pivot_longer把宽格式数据转为长格式,再按组统计唯一值:

result2 <- df %>%
  tidyr::pivot_longer(cols = DX1:DX4, names_to = "dx_col", values_to = "dx_value") %>%
  group_by(id) %>%
  summarize(unique_dx_count = n_distinct(dx_value, na.rm = TRUE))

print(result2)

这个方法的结果和方法一完全一致,适合需要对转长后的数据做更多后续处理的场景。

内容的提问来源于stack exchange,提问作者mh3177

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 15:27:11