You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中基于2022年的ID筛选并移除对应所有行?

移除2022年出现过的ID的所有行(tidyverse实现)

原始数据集

生成代码

df <- data.frame(ID = c(11, 12, 11, 12, 13, 11, 13, 14, 12, 14, 15),
       year = c(2022, 2022, 2021, 2021, 2021, 2020, 2020, 2020, 2019, 2019, 2019),
       name = c("Jay", "Rahul", "Jay", "Rahul", "Pierre", "Jay", "Pierre",
                "Shannon", "Rahul", "Shannon", "Zachary"))

数据表格

IDyearname
112022Jay
122022Rahul
112021Jay
122021Rahul
132021Pierre
112020Jay
132020Pierre
142020Shannon
122019Rahul
142019Shannon
152019Zachary

需求说明

移除所有在2022年出现过的ID对应的全部行,数据集规模较大(约9000行、70个年份),需使用tidyverse、dplyr包高效实现。

解决方案

提供两种高效实现方式,均基于tidyverse生态:

方法1:分组后过滤(单管道操作)

利用group_by按ID分组,通过any(year == 2022)判断该ID是否存在2022年记录,反向保留无2022年记录的ID的所有行:

library(tidyverse)

result_df <- df %>%
  group_by(ID) %>%
  filter(!any(year == 2022)) %>%
  ungroup()

方法2:先提取排除ID列表再过滤

先提取2022年出现的所有唯一ID,再过滤掉这些ID的所有行,适合需要单独查看排除ID的场景:

library(tidyverse)

# 获取2022年出现的所有唯一ID
exclude_ids <- df %>%
  filter(year == 2022) %>%
  pull(ID) %>%
  unique()

# 过滤掉排除ID的所有行
result_df <- df %>%
  filter(!ID %in% exclude_ids)

输出结果

IDyearname
132021Pierre
132020Pierre
142020Shannon
142019Shannon
152019Zachary

内容的提问来源于stack exchange,提问作者1shh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 15:57:49