You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言合并数据框中start与end列对应连续范围的相邻行

基因组连续区间合并R实现方案

完整可运行代码

# 加载依赖包,如未安装先运行 install.packages("dplyr")
library(dplyr)

# 示例测试数据
genomic_ranges <- data.frame(
  sample_ID = c("A", "B", "B", "B", "C"),
  start = c(1, 20, 30, 40, 250),
  end =  c(5, 30, 40, 70, 400),
  feature = c("normal", "DUP", "DUP", "DUP", "DUP")
)

# 核心合并逻辑
merged_ranges <- genomic_ranges %>%
  # 先按分组字段+起始位置升序排序,确保区间顺序正确
  arrange(sample_ID, feature, start) %>%
  # 按样本、变异类型分组,仅组内执行合并逻辑
  group_by(sample_ID, feature) %>%
  # 标记连续区间块:当前行start不等于上一行end时,新开一个块
  mutate(block_id = cumsum(start != lag(end, default = first(start) - 1))) %>%
  # 按样本、变异类型、区间块二次分组
  group_by(sample_ID, feature, block_id) %>%
  # 汇总每个连续块的最小起始、最大终止位置
  summarise(
    start = min(start),
    end = max(end),
    .groups = "drop"
  ) %>%
  # 移除辅助计算的block_id字段
  select(-block_id)

运行结果

> merged_ranges
# A tibble: 3 × 4
  sample_ID feature start   end
  <chr>     <chr>   <dbl> <dbl>
1 A         normal      1     5
2 B         DUP        20    70
3 C         DUP       250   400

注意事项

  • 代码针对20万行规模的数据运行效率足够,无需使用低效率的for循环实现
  • 若需适配存在重叠(不止是首尾相等)的区间合并场景,只需把判断条件start != lag(end)改为start > lag(end)即可
  • 运行前请确保你的数据中start、end列均为整数类型,避免字符型比较出现误差

内容的提问来源于stack exchange,提问作者OncoComp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 11:45:05