如何在R语言中跳过指定分组最大索引后的指定行数
解决方案
可以通过识别连续的class分组,定位目标分组的末尾行,标记后续两行并删除,具体实现如下:
方法一:使用dplyr包
library(dplyr) # 生成连续分组ID,标记分组内的最后一行及目标分组 dd_processed <- dd %>% mutate(group_id = data.table::rleid(class)) %>% group_by(group_id) %>% mutate(is_last = row_number() == n(), target_group = class %in% c(2, 4)) %>% ungroup() %>% mutate(to_remove = FALSE) # 定位目标分组的最后一行索引,标记后续两行需要删除 target_last_indices <- which(dd_processed$is_last & dd_processed$target_group) for(idx in target_last_indices){ if(idx + 1 <= nrow(dd_processed)) dd_processed$to_remove[idx + 1] <- TRUE if(idx + 2 <= nrow(dd_processed)) dd_processed$to_remove[idx + 2] <- TRUE } # 过滤得到结果 result <- dd_processed %>% filter(!to_remove) %>% select(rate, class) print(result)
方法二:使用data.table包(高效处理大数据集)
library(data.table) setDT(dd) # 生成连续分组ID,提取每个分组的起止索引和class dd[, group_id := rleid(class)] group_info <- dd[, .(start = .I[1], end = .I[.N], class = first(class)), by = group_id] # 获取目标分组的末尾索引,生成需要删除的行索引 target_ends <- group_info[class %in% c(2,4), end] remove_indices <- unlist(lapply(target_ends, function(x) c(x+1, x+2))) remove_indices <- remove_indices[remove_indices <= nrow(dd)] # 删除指定行得到结果 result <- dd[-remove_indices] print(result)
结果验证
运行上述代码后,输出结果与期望一致:
rate class 1 0.5 9 2 0.7 9 3 0.6 4 4 0.5 4 5 0.3 4 6 0.9 4 7 0.8 1 8 0.6 1 9 0.2 4 10 0.1 4 11 0.2 3 12 1.0 2 13 0.7 2 14 0.6 9
内容的提问来源于stack exchange,提问作者monckeyyL
相关产品推荐
相关产品推荐

