You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按ID筛选满足至少2个指定条件的记录并添加index列

按ID分组判断多条件覆盖情况的R语言实现

原始数据

df <- read.table(header=TRUE, text="id code
                1 A
                1 B
                1 C
                2 A
                2 A
                2 A
                3 A
                3 B
                3 A")

需求说明

按id分组,判断每个分组是否覆盖了至少2个指定条件(条件为code等于"A"、"B"、"C"中的至少两种),并新增index列:满足条件的分组内所有行index为1,否则为0。期望输出如下:

df_output <- read.table(header=TRUE, text="id code index
                1 A 1
                1 B 1
                1 C 1
                2 A 0
                2 A 0
                2 A 0
                3 A 1
                3 B 1
                3 A 1")

原代码问题

你之前的代码仅判断每行code是否匹配单个条件,没有统计分组内满足的不同条件数量,无法实现“至少2个条件”的阈值判断:

df_output = df %>% 
     group_by(id) %>%
     mutate(index = ifelse(grepl(conditionA|conditionB|conditionC, code), 1, 0))

正确实现代码

方法一(分步清晰版)

library(dplyr)

# 定义目标条件集合
target_codes <- c("A", "B", "C")

df_output <- df %>%
  group_by(id) %>%
  # 统计分组内属于目标集合的唯一code数量
  mutate(unique_target_count = n_distinct(intersect(code, target_codes))) %>%
  # 根据数量判断是否满足条件,生成index列
  mutate(index = ifelse(unique_target_count >= 2, 1, 0)) %>%
  # 移除中间辅助列(可选操作)
  select(-unique_target_count) %>%
  ungroup()

方法二(简洁版)

library(dplyr)

df_output <- df %>%
  group_by(id) %>%
  # 直接统计目标条件的唯一覆盖数,转成1/0格式
  mutate(index = as.integer(n_distinct(intersect(code, c("A", "B", "C"))) >= 2)) %>%
  ungroup()

逻辑说明

  • intersect(code, target_codes):提取当前分组中属于目标条件的code值
  • n_distinct(...):统计上述结果中的唯一值数量,即该分组覆盖的不同条件数
  • as.integer(...):将逻辑判断结果(TRUE/FALSE)转换为1/0,等价于ifelse的作用

内容的提问来源于stack exchange,提问作者Economist_Ayahuasca

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 16:27:25