如何用dplyr按组筛选value列中的共同值
用dplyr按组筛选保留不同id组的共同value值
示例数据
library(dplyr) df <- structure(list(id = c(1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L, 2L), value = c(1, 2, 3, 5, 6, 6, 3, 2, 0, 10)), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -10L))
需求说明
需要筛选出所有id分组中都存在的value值,保留这些value对应的所有行,最终得到如下结果:
# A tibble: 6 × 2 id value <int> <dbl> 1 1 2 2 1 3 3 1 6 4 2 2 5 2 3 6 2 6
解决方案
核心思路是先找出所有分组共有的value集合,再基于这个集合筛选原数据:
方法一:分步实现
# 1. 提取所有id组共同的value值 common_values <- df %>% group_by(value) %>% summarise(出现的组数量 = n_distinct(id)) %>% filter(出现的组数量 == n_distinct(df$id)) %>% pull(value) # 2. 筛选原数据中符合条件的行 result <- df %>% filter(value %in% common_values) %>% arrange(id, value) # 排序以匹配期望输出,可选 result
方法二:链式合并(更简洁)
df %>% filter(value %in% ( df %>% group_by(value) %>% summarise(n_groups = n_distinct(id)) %>% filter(n_groups == n_distinct(df$id)) %>% pull(value) )) %>% arrange(id, value)
代码解释
- 分组统计阶段:按
value分组后,用n_distinct(id)统计每个值出现在多少个不同的id组里,再筛选出出现次数等于总组数的value,这些就是所有组的共同值。 - 筛选阶段:用
value %in% common_values保留原数据中属于共同值的行,最后用arrange排序让结果和期望输出一致。
内容的提问来源于stack exchange,提问作者HoelR
相关产品推荐
相关产品推荐

