You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于DataFrame特定列条件的行筛选:保留Treat=1起始的个体数据

按ID保留Treat=1之后所有行的解决方案

需求说明

处理随时间t重复记录的动物ID数据:

  • 保留每个Id从首次出现Treat=1的行开始的所有后续行
  • 若某个Id没有Treat=1的记录,则丢弃该ID的所有行

示例数据

bd = data.frame(
  t = c(1,2,3,4,5,6,1,2,3,4,5,6,1,2,3,4,5,6),
  Id = rep(1:3, each=6), # 修正原数据的Id生成逻辑,确保3个个体各6条记录
  Treat = c(0,0,1,0,0,0, 0,0,0,0,1,0,0,0,0,0,0,0)
)

方法一:使用dplyr包(推荐)

library(dplyr)

result <- bd %>%
  group_by(Id) %>%
  # 标记当前ID首次出现Treat=1的行号,无则设为无穷大
  mutate(first_treat = ifelse(any(Treat == 1), min(which(Treat == 1)), Inf)) %>%
  # 保留首次处理及之后的所有行
  filter(row_number() >= first_treat) %>%
  # 移除辅助计算列
  select(-first_treat) %>%
  ungroup()

# 查看结果
print(result)

方法二:Base R实现

# 计算每个ID首次出现Treat=1的位置
first_treat_pos <- tapply(bd$Treat, bd$Id, function(x) {
  pos <- which(x == 1)
  if (length(pos) == 0) Inf else min(pos)
})

# 为每行匹配对应的首次处理位置
bd$first_treat <- first_treat_pos[as.character(bd$Id)]

# 过滤目标行并移除辅助列
result_base <- bd[seq(nrow(bd)) >= bd$first_treat, ]
result_base <- result_base[, !names(result_base) %in% "first_treat"]

# 查看结果
print(result_base)

结果说明

处理后的数据会保留:

  • Id=1的t=3至t=6行
  • Id=2的t=5至t=6行
  • Id=3的所有行被丢弃

内容的提问来源于stack exchange,提问作者Buczinski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 19:20:31