R语言筛选分组ID下某列仅含唯一值的所有数据行
鸟类观测数据固定位置个体筛选方法
需求描述
现有鸟类观测数据框df,包含3个字段:
Bird.ID:鸟类个体唯一编号Location:观测点位Date:观测日期
需要完成子集筛选:仅保留所有观测记录中始终出现在同一位置的鸟类个体对应的全部观测行,只要某只个体在2个及以上不同点位被观测到,就剔除该个体的所有记录。
样例输入
df Bird.ID Location Date 1 Plot1 22/02/2022 2 Plot5 22/02/2022 3 Plot1 22/02/2022 1 Plot1 24/02/2022 1 Plot1 26/02/2022 1 Plot1 22/03/2022 2 Plot5 22/03/2022 2 Plot5 14/04/2022 3 Plot2 14/04/2022 3 Plot3 22/06/2022
样例中编号为3的个体先后在Plot1、Plot2、Plot3三个点位出现,属于移动个体,需要全部剔除;保留始终在Plot1的1号个体、始终在Plot5的2号个体的所有记录。
预期输出
output Bird.ID Location Date 1 Plot1 22/02/2022 2 Plot5 22/02/2022 1 Plot1 24/02/2022 1 Plot1 26/02/2022 1 Plot1 22/03/2022 2 Plot5 22/03/2022 2 Plot5 14/04/2022
实现代码
方法1:dplyr 实现(代码简洁易读)
library(dplyr) output <- df %>% group_by(Bird.ID) %>% # 按个体分组后,仅保留唯一观测点数量为1的分组所有行 filter(n_distinct(Location) == 1) %>% ungroup()
方法2:R基础语法实现(无需加载第三方包)
# 先筛选出所有仅在单个点位出现的鸟类ID valid_birds <- names(which( tapply(df$Location, df$Bird.ID, function(x) length(unique(x)) == 1) )) # 匹配原数据完成子集提取 output <- df[df$Bird.ID %in% valid_birds, ]
两种方法运行后得到的结果和预期输出完全一致。
内容的提问来源于stack exchange,提问作者Andre230
相关产品推荐
相关产品推荐

