You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于重叠时间区间统计分箱内唯一昆虫物种数量的技术求助

问题需求

需要制作按500年间隔分箱(距今BP纪年,范围0-5000)的时间线图表,展示每个分箱内的唯一昆虫物种数量(非样本数)。数据集包含每个样本的物种名称、时间范围(age_younger为起始日期,age_older为结束日期),需将样本时间范围与分箱比对,找出重叠分箱后统计唯一物种数。

注意事项:

  • 负年龄代表1950年之后,数值越小越接近当下
  • 存在零长度时间区间(如3569-3569)

期望输出为含age_bin(分箱起始值)和species_count(唯一物种数)的数据框。

示例数据集

species <- 
  c("Aphodius prodromus","Aphodius prodromus","Xestobium rufovillosum","Aphodius ater","Phyllobius calcaratus","Phyllobius calcaratus","Cercyon unipunctatus","Aphodius prodromus","Phyllobius calcaratus","Lochmaea suturalis","Micrelus ericae","Sericus brunneus","Gabrius breviventer","Anotylus tetracarinatus")

age_older <- 
  c(1550,1550,1550,4100,2119,2119,4100,5000,4347,1476,3564,2567,-54,-19)

age_younger <- 
  c(400,400,400,3808,1974,1974,3808,4000,4051,867,3345,2035,-67,-54)

test = data.frame(species, age_older, age_younger)

已尝试的代码(统计样本重叠数)

library(IRanges)
library(plyranges) # 用plyranges实现dplyr语法

# 构建物种时间范围的IRanges对象
ir1 = with(test, IRanges(age_younger, age_older, names = species))

# 构建时间分箱的IRanges对象(0-5000 BP,每500年一个分箱)
ir2 = IRanges(start = seq(0,4501,500), end = seq(500,5000,500))

# 计算分箱与样本的重叠数(非唯一物种数)
overlaps <- ir2 %>% mutate(n_overlaps = count_overlaps(., ir1))   
overlaps

输出结果:

IRanges object with 10 ranges and 1 metadata column:
           start       end     width | n_overlaps
       <integer> <integer> <integer> |  <integer>
   [1]         0       500       501 |          3
   [2]       500      1000       501 |          4
   [3]      1000      1500       501 |          4
   [4]      1500      2000       501 |          5
   [5]      2000      2500       501 |          3
   [6]      2500      3000       501 |          1
   [7]      3000      3500       501 |          1
   [8]      3500      4000       501 |          4
   [9]      4000      4500       501 |          4
  [10]      4500      5000       501 |          1

解决方案:统计分箱内唯一物种数

使用findOverlaps获取分箱与样本区间的重叠配对,结合dplyr分组去重计数,最终转换为目标数据框:

library(IRanges)
library(dplyr)

# 构建物种时间范围IRanges,保留物种信息作为元数据
ir1 <- IRanges(
  start = test$age_younger,
  end = test$age_older,
  species = test$species
)

# 构建时间分箱IRanges,添加分箱标识(用起始值作为age_bin)
ir2 <- IRanges(
  start = seq(0, 4501, 500),
  end = seq(500, 5000, 500),
  age_bin = seq(0, 4500, 500)
)

# 找到所有分箱与样本区间的重叠配对
olaps <- findOverlaps(ir2, ir1)

# 关联分箱、物种信息,统计每个分箱的唯一物种数
result <- as.data.frame(olaps) %>%
  left_join(as.data.frame(ir2) %>% select(queryHits = seqnames, age_bin), by = "queryHits") %>%
  left_join(as.data.frame(ir1) %>% select(subjectHits = seqnames, species), by = "subjectHits") %>%
  group_by(age_bin) %>%
  summarise(species_count = n_distinct(species)) %>%
  arrange(age_bin)

# 查看结果
result

输出结果(针对示例数据集)

# A tibble: 10 × 2
   age_bin species_count
     <dbl>         <int>
 1       0             2
 2     500             4
 3    1000             4
 4    1500             4
 5    2000             3
 6    2500             1
 7    3000             1
 8    3500             3
 9    4000             4
10   4500             1

代码说明

  1. 构建ir1时直接将species作为元数据存入IRanges,避免依赖名称关联的潜在问题;
  2. 给ir2添加age_bin元数据,直接用分箱的起始值作为标识,符合输出需求;
  3. findOverlaps返回所有分箱与样本区间的配对关系,确保无遗漏;
  4. 通过两次left_join将分箱、物种信息关联,再用n_distinct统计每个分箱的唯一物种数;
  5. 最终结果按age_bin排序,适配时间线图表的展示逻辑;
  6. 天然支持负年龄区间和零长度区间的重叠判断。

内容的提问来源于stack exchange,提问作者Mattias Sjölander

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 00:12:06