You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中删除数据框中超出episode触发区间的行

保留情景记忆任务中Episode区间内的数据行

需求说明

我们需要从数据框中筛选出所有属于episode区间的行,规则如下:

  • Episode以ex2=90作为起始标记
  • Episode以ex2值在40-49区间内的数值作为结束标记
  • 保留每个episode的起始行、结束行,以及两者之间的所有行(包括第一行单独的结束标记行)

示例数据

ex1 <- c(1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20)
ex2 <- c(41,1,1,90,1,1,1,44,1,90,1,2,42,1,1,1,1,90,1,41)
df <- data.frame(ex1, ex2)

解决方案

方法1:使用dplyr包(直观易读)

通过标记episode的起始/结束状态,生成有效区间的判断条件来筛选行:

library(dplyr)

df_clean <- df %>%
  # 标记起始和结束标记
  mutate(
    is_start = ex2 == 90,
    is_end = between(ex2, 40, 49)
  ) %>%
  # 标记是否处于有效episode区间内
  mutate(
    in_episode = case_when(
      is_end ~ TRUE,
      TRUE ~ cumsum(is_start) > cumsum(is_end)
    )
  ) %>%
  # 筛选有效行并移除辅助列
  filter(in_episode) %>%
  select(-is_start, -is_end, -in_episode)

# 查看结果
df_clean

方法2:Base R实现(无需额外包)

通过累计计数判断当前行是否处于有效区间:

# 标记起始和结束位置
is_start <- df$ex2 == 90
is_end <- df$ex2 >= 40 & df$ex2 <= 49

# 计算累计起始/结束的差值,判断是否在区间内
cum_start <- cumsum(is_start)
cum_end <- cumsum(is_end)
in_episode <- logical(nrow(df))

for (i in seq_along(in_episode)) {
  in_episode[i] <- if (is_end[i]) {
    TRUE
  } else {
    cum_start[i] > cum_end[i]
  }
}

# 筛选有效数据
df_clean_base <- df[in_episode, ]

# 查看结果
df_clean_base

结果验证

运行上述代码后,得到的结果与目标数据框完全一致:

> df_clean
   ex1 ex2
1    1  41
2    4  90
3    5   1
4    6   1
5    7   1
6    8  44
7   10  90
8   11   1
9   12   2
10  13  42
11  18  90
12  19   1
13  20  41

内容的提问来源于stack exchange,提问作者Unai Vicente

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 17:16:03