基于ID与Group组合创建Indicator列的R语言实现求助
解决方案
可以用dplyr包的分组功能快速实现这个需求,核心思路是按ID分组,针对每个组的Group值集合生成对应的Indicator:
步骤1:构造原始数据集
library(dplyr) # 你的原始数据 df <- tibble( ID = c(10, 11, 13, 13, 15, 15, 17, 17, 17, 19, 19, 23, 23, 23), Group = c("Red", "Red", "Blue", "Red", "Blue", "Blue", "Blue", "Red", "Red", "Blue", "Undecided", "Blue", "Undecided", "Undecided") )
步骤2:生成Indicator列
df_processed <- df %>% group_by(ID) %>% mutate( # 提取当前ID下所有不重复的Group值,保留首次出现的顺序 unique_groups = list(unique(Group)), Indicator = case_when( # 情况1:该ID只有1个Group值且仅出现1次 → 标记为Single length(unique_groups[[1]]) == 1 & n() == 1 ~ "Single", # 情况2:该ID只有1个Group值但出现多次 → 格式为"Group-Group" length(unique_groups[[1]]) == 1 ~ paste(unique_groups[[1]], unique_groups[[1]], sep = "-"), # 情况3:该ID有多个不同Group值 → 用"-"连接所有不重复的Group值 TRUE ~ paste(unique_groups[[1]], collapse = "-") ) ) %>% # 移除临时辅助列 select(-unique_groups) %>% ungroup()
处理结果
运行后得到的df_processed如下:
# A tibble: 14 × 3 ID Group Indicator <dbl> <chr> <chr> 1 10 Red Single 2 11 Red Single 3 13 Blue Blue-Red 4 13 Red Blue-Red 5 15 Blue Blue-Blue 6 15 Blue Blue-Blue 7 17 Blue Blue-Red 8 17 Red Blue-Red 9 17 Red Blue-Red 10 19 Blue Blue-Undecided 11 19 Undecided Blue-Undecided 12 23 Blue Blue-Undecided 13 23 Undecided Blue-Undecided 14 23 Undecided Blue-Undecided
说明
- 如果希望按字母顺序排列Group值(而非首次出现顺序),只需把
unique(Group)改成sort(unique(Group))即可; - 整个逻辑通过分组处理避免了复杂的
case_when嵌套,每个ID的规则判断只在组内执行一次,效率更高。
内容的提问来源于stack exchange,提问作者Sundown Brownbear
相关产品推荐
相关产品推荐

