You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中按条件分组提取DataFrame行的实现方案

R实现按ID分组筛选指定行

示例数据构造

先还原你提供的DataFrame:

df <- data.frame(
  ID = c(1,2,2,2,3,3,4,4,4,4,5,5,5),
  Decision = c("yes","no","yes","no","no","no","no","no","yes","no","no","no","no"),
  Time = as.POSIXct(c(
    "2017-06-25 17:30:30",
    "2017-06-15 17:32:30",
    "2017-06-15 17:30:30",
    "2017-06-15 17:20:30",
    "2017-06-22 17:30:30",
    "2017-06-21 18:31:30",
    "2017-05-25 17:30:30",
    "2017-05-06 18:30:30",
    "2017-04-06 19:30:30",
    "2017-04-07 13:30:30",
    "2018-06-25 18:30:30",
    "2018-06-25 19:30:30",
    "2018-06-25 21:30:30"
  ))
)

解决方案1:使用dplyr(tidyverse生态)

这是最简洁直观的实现方式,适合日常数据处理:

library(dplyr)

result <- df %>%
  group_by(ID) %>%
  filter(
    # 组内存在yes时,保留所有yes行(若单个ID有多个yes可按需调整,你的示例中每个ID仅一个yes)
    Decision == "yes" |
    # 组内全为no时,保留时间最晚的行
    (all(Decision == "no") & Time == max(Time))
  ) %>%
  ungroup()

# 输出结果
print(result)

解决方案2:Base R实现

无需额外安装包,适合轻量场景:

result_base <- do.call(rbind, lapply(split(df, df$ID), function(group) {
  if (any(group$Decision == "yes")) {
    subset(group, Decision == "yes")
  } else {
    subset(group, Time == max(group$Time))
  }
}))

# 重置行名避免混乱
rownames(result_base) <- NULL
print(result_base)

输出结果

两种方法都会得到你期望的DataFrame:

IDDecisionTime
1yes2017-06-25 17:30:30
2yes2017-06-15 17:30:30
3no2017-06-22 17:30:30
4yes2017-04-06 19:30:30
5no2018-06-25 21:30:30

内容的提问来源于stack exchange,提问作者cristi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 11:11:16