在R语言中按条件分组提取DataFrame行的实现方案
R实现按ID分组筛选指定行
示例数据构造
先还原你提供的DataFrame:
df <- data.frame( ID = c(1,2,2,2,3,3,4,4,4,4,5,5,5), Decision = c("yes","no","yes","no","no","no","no","no","yes","no","no","no","no"), Time = as.POSIXct(c( "2017-06-25 17:30:30", "2017-06-15 17:32:30", "2017-06-15 17:30:30", "2017-06-15 17:20:30", "2017-06-22 17:30:30", "2017-06-21 18:31:30", "2017-05-25 17:30:30", "2017-05-06 18:30:30", "2017-04-06 19:30:30", "2017-04-07 13:30:30", "2018-06-25 18:30:30", "2018-06-25 19:30:30", "2018-06-25 21:30:30" )) )
解决方案1:使用dplyr(tidyverse生态)
这是最简洁直观的实现方式,适合日常数据处理:
library(dplyr) result <- df %>% group_by(ID) %>% filter( # 组内存在yes时,保留所有yes行(若单个ID有多个yes可按需调整,你的示例中每个ID仅一个yes) Decision == "yes" | # 组内全为no时,保留时间最晚的行 (all(Decision == "no") & Time == max(Time)) ) %>% ungroup() # 输出结果 print(result)
解决方案2:Base R实现
无需额外安装包,适合轻量场景:
result_base <- do.call(rbind, lapply(split(df, df$ID), function(group) { if (any(group$Decision == "yes")) { subset(group, Decision == "yes") } else { subset(group, Time == max(group$Time)) } })) # 重置行名避免混乱 rownames(result_base) <- NULL print(result_base)
输出结果
两种方法都会得到你期望的DataFrame:
| ID | Decision | Time |
|---|---|---|
| 1 | yes | 2017-06-25 17:30:30 |
| 2 | yes | 2017-06-15 17:30:30 |
| 3 | no | 2017-06-22 17:30:30 |
| 4 | yes | 2017-04-06 19:30:30 |
| 5 | no | 2018-06-25 21:30:30 |
内容的提问来源于stack exchange,提问作者cristi
相关产品推荐
相关产品推荐

