如何基于多列筛选DataFrame?保留同组内origin与destination一致的数据
解决方案
需求:按ID和planes分组,仅保留组内所有行的origin和destination完全一致的分组数据。
方法一:使用dplyr包
这是tidyverse生态下的常用方法,代码简洁易读:
library(dplyr) # 原始数据 data <- data.frame( ID = c(111, 111, 111,111, 333, 333, 333,333), planes = c(2, 2, 3, 3, 4, 4, 5, 5), origin = c(3, 3, 6, 8, 5, 5, 7, 9), destination = c(9, 9, 10, 20, 11, 11, 13, 25) ) # 分组筛选 result <- data %>% group_by(ID, planes) %>% # 筛选组内origin和destination都只有唯一值的分组 filter(n_distinct(origin) == 1 & n_distinct(destination) == 1) %>% ungroup() print(result)
运行后输出:
# A tibble: 4 × 4 ID planes origin destination <dbl> <dbl> <dbl> <dbl> 1 111 2 3 9 2 111 2 3 9 3 333 4 5 11 4 333 4 5 11
方法二:Base R实现
无需加载额外包,使用ave()函数完成分组判断:
# 原始数据 data <- data.frame( ID = c(111, 111, 111,111, 333, 333, 333,333), planes = c(2, 2, 3, 3, 4, 4, 5, 5), origin = c(3, 3, 6, 8, 5, 5, 7, 9), destination = c(9, 9, 10, 20, 11, 11, 13, 25) ) # 标记需要保留的行 data$keep <- with(data, ave(origin, ID, planes, FUN = function(x) length(unique(x)) == 1) & ave(destination, ID, planes, FUN = function(x) length(unique(x)) == 1)) # 筛选并移除标记列 result_base <- data[data$keep == 1, !names(data) %in% "keep"] print(result_base)
运行后输出与方法一一致。
内容的提问来源于stack exchange,提问作者Xaviermoros
相关产品推荐
相关产品推荐

