You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用tidyverse动态保留重复ID的Q40=Plan行并剔除对应NoPlan行?

使用tidyverse处理重复参与者数据的解决方案

首先加载所需包并定义原始数据集:

library(tidyverse)

# 原始数据集
df <- structure(
  list(
    ID = c("Joseph", "Cyntia", "Paul", "Ana", "Paul", "Maria", "Ana"),
    Q40 = c("Plan", "NoPlan", "Plan", "Plan", "NoPlan", "Plan", "NoPlan")
  ),
  row.names = c(NA, -7L),
  class = c("tbl_df", "tbl", "data.frame")
)

核心处理代码

df_cleaned <- df %>%
  # 按ID分组,标记该ID是否同时存在Plan和NoPlan两种记录
  group_by(ID) %>%
  mutate(has_conflict = all(c("Plan", "NoPlan") %in% Q40)) %>%
  ungroup() %>%
  # 过滤规则:无冲突ID保留所有行;有冲突ID仅保留Q40=Plan的行
  filter(!has_conflict | Q40 == "Plan") %>%
  # 移除辅助标记列
  select(-has_conflict)

代码说明

  • 标记冲突ID:通过group_by(ID)分组后,用all(c("Plan", "NoPlan") %in% Q40)判断每个ID是否同时包含两种记录,生成has_conflict辅助列,精准定位需要处理的重复参与者。
  • 精准过滤:filter(!has_conflict | Q40 == "Plan")完全满足你的要求:
    • 不会过滤所有NoPlan行(比如Cyntia的NoPlan记录会保留)
    • 不是简单保留首行,而是基于Q40的值筛选
    • 没有使用全局distinct,仅针对有冲突的ID做处理

最终输出

运行代码后得到的清理后数据集:

df_cleaned
#> # A tibble: 5 × 2
#>   ID      Q40  
#>   <chr>   <chr>
#> 1 Joseph  Plan 
#> 2 Cyntia  NoPlan
#> 3 Paul    Plan 
#> 4 Ana     Plan 
#> 5 Maria   Plan

内容的提问来源于stack exchange,提问作者Larissa Cury

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 06:37:09