在R语言中删除数据框中超出episode触发区间的行
保留情景记忆任务中Episode区间内的数据行
需求说明
我们需要从数据框中筛选出所有属于episode区间的行,规则如下:
- Episode以
ex2=90作为起始标记 - Episode以
ex2值在40-49区间内的数值作为结束标记 - 保留每个episode的起始行、结束行,以及两者之间的所有行(包括第一行单独的结束标记行)
示例数据
ex1 <- c(1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20) ex2 <- c(41,1,1,90,1,1,1,44,1,90,1,2,42,1,1,1,1,90,1,41) df <- data.frame(ex1, ex2)
解决方案
方法1:使用dplyr包(直观易读)
通过标记episode的起始/结束状态,生成有效区间的判断条件来筛选行:
library(dplyr) df_clean <- df %>% # 标记起始和结束标记 mutate( is_start = ex2 == 90, is_end = between(ex2, 40, 49) ) %>% # 标记是否处于有效episode区间内 mutate( in_episode = case_when( is_end ~ TRUE, TRUE ~ cumsum(is_start) > cumsum(is_end) ) ) %>% # 筛选有效行并移除辅助列 filter(in_episode) %>% select(-is_start, -is_end, -in_episode) # 查看结果 df_clean
方法2:Base R实现(无需额外包)
通过累计计数判断当前行是否处于有效区间:
# 标记起始和结束位置 is_start <- df$ex2 == 90 is_end <- df$ex2 >= 40 & df$ex2 <= 49 # 计算累计起始/结束的差值,判断是否在区间内 cum_start <- cumsum(is_start) cum_end <- cumsum(is_end) in_episode <- logical(nrow(df)) for (i in seq_along(in_episode)) { in_episode[i] <- if (is_end[i]) { TRUE } else { cum_start[i] > cum_end[i] } } # 筛选有效数据 df_clean_base <- df[in_episode, ] # 查看结果 df_clean_base
结果验证
运行上述代码后,得到的结果与目标数据框完全一致:
> df_clean ex1 ex2 1 1 41 2 4 90 3 5 1 4 6 1 5 7 1 6 8 44 7 10 90 8 11 1 9 12 2 10 13 42 11 18 90 12 19 1 13 20 41
内容的提问来源于stack exchange,提问作者Unai Vicente
相关产品推荐
相关产品推荐

