You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中使用dplyr的group_by多列统计时如何消除重复行?

问题分析与解决

先看你的原始数据:

library(dplyr)
mydat <- data.frame(ID = c(123, 123, 111, 111, 111), 
           class = c("A", "A", "A", "A", "B"),
           new_ID = c(999, 872, 999, 1566, 254))
mydat
#    ID class new_ID
# 1 123     A    999
# 2 123     A    872
# 3 111     A    999
# 4 111     A   1566
# 5 111     B    254

你用mutate出现重复行的原因很直接:mutate的作用是给原数据的每一行添加新列,哪怕是分组计算,它也会把分组统计的结果复制到该组的每一行里,所以原数据有多少行,结果就有多少行。

要得到每个ID-class配对的唯一统计行,你需要用summarise()(或美式拼写summarize()),它会把每个分组聚合为单独一行:

mydat %>% 
  group_by(ID, class) %>% 
  summarise(n_new_ID = n_distinct(new_ID), .groups = "keep") %>% 
  select(ID, class, n_new_ID)

执行结果就是你要的:

# A tibble: 3 × 3
# Groups:   ID, class [3]
     ID class n_new_ID
  <dbl> <chr>    <int>
1   123 A            2
2   111 A            2
3   111 B            1

另外,用n_distinct(new_ID)比length(unique(new_ID))更符合dplyr的风格,效果完全一致。.groups = "keep"用来保留分组信息,如果你不需要分组结构,可以改成.groups = "drop"。

内容的提问来源于stack exchange,提问作者Adrian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 19:15:43