基于规则条件标记R语言数据集ID的实现方案问询
R语言数据集的条件式标记实现方案
样本数据集
首先定义示例数据集:
data <- data.frame(id = c(1,1,1,1,1,1, 2,2,2, 3,3,3), cat1 = c("A","A","A","B","B","B", "A","A","A", "A","A","B"), levels = c("L1","L3","L4","L2","L1","L3", "L1","L2","L2", "L1","L2","L1"))
样本数据预览:
id cat1 levels 1 1 A L1 2 1 A L3 3 1 A L4 4 1 B L2 5 1 B L1 6 1 B L3 7 2 A L1 8 2 A L2 9 2 A L2 10 3 A L1 11 3 A L2 12 3 B L1
标记规则
需按以下规则为每个id统一标记label:
- 规则a:若该
id的cat1="A"对应的levels包含L3或L4,且存在cat1="B",标记为Rule_satisfied - 规则b:若该
id的cat1="A"对应的levels仅包含L1或L2,且不存在cat1="B",标记为Rule_NotSatisfied - 规则c:若该
id的cat1="A"对应的levels仅包含L1或L2,但存在cat1="B",标记为Rule_violation
实现代码
使用dplyr包进行分组逻辑判断,步骤清晰且高效:
# 首次使用需先安装dplyr:install.packages("dplyr") library(dplyr) data.1 <- data %>% # 按id分组,确保同一id的标记统一 group_by(id) %>% mutate( # 计算两个核心判断条件 has_high_level = any(cat1 == "A" & levels %in% c("L3", "L4")), has_B = any(cat1 == "B"), # 根据规则匹配生成label label = case_when( has_high_level & has_B ~ "Rule_satisfied", !has_high_level & !has_B ~ "Rule_NotSatisfied", !has_high_level & has_B ~ "Rule_violation", TRUE ~ "Unknown" # 兜底逻辑,实际不会触发 ) ) %>% # 取消分组,恢复原始数据结构 ungroup() %>% # 移除中间辅助计算列(可选操作) select(-has_high_level, -has_B)
输出结果
运行上述代码后,得到目标数据集:
id cat1 levels label 1 1 A L1 Rule_satisfied 2 1 A L3 Rule_satisfied 3 1 A L4 Rule_satisfied 4 1 B L2 Rule_satisfied 5 1 B L1 Rule_satisfied 6 1 B L3 Rule_satisfied 7 2 A L1 Rule_NotSatisfied 8 2 A L2 Rule_NotSatisfied 9 2 A L2 Rule_NotSatisfied 10 3 A L1 Rule_violation 11 3 A L2 Rule_violation 12 3 B L1 Rule_violation
内容的提问来源于stack exchange,提问作者amisos55
相关产品推荐
相关产品推荐

