基于重叠时间区间统计分箱内唯一昆虫物种数量的技术求助
问题需求
需要制作按500年间隔分箱(距今BP纪年,范围0-5000)的时间线图表,展示每个分箱内的唯一昆虫物种数量(非样本数)。数据集包含每个样本的物种名称、时间范围(age_younger为起始日期,age_older为结束日期),需将样本时间范围与分箱比对,找出重叠分箱后统计唯一物种数。
注意事项:
- 负年龄代表1950年之后,数值越小越接近当下
- 存在零长度时间区间(如3569-3569)
期望输出为含age_bin(分箱起始值)和species_count(唯一物种数)的数据框。
示例数据集
species <- c("Aphodius prodromus","Aphodius prodromus","Xestobium rufovillosum","Aphodius ater","Phyllobius calcaratus","Phyllobius calcaratus","Cercyon unipunctatus","Aphodius prodromus","Phyllobius calcaratus","Lochmaea suturalis","Micrelus ericae","Sericus brunneus","Gabrius breviventer","Anotylus tetracarinatus") age_older <- c(1550,1550,1550,4100,2119,2119,4100,5000,4347,1476,3564,2567,-54,-19) age_younger <- c(400,400,400,3808,1974,1974,3808,4000,4051,867,3345,2035,-67,-54) test = data.frame(species, age_older, age_younger)
已尝试的代码(统计样本重叠数)
library(IRanges) library(plyranges) # 用plyranges实现dplyr语法 # 构建物种时间范围的IRanges对象 ir1 = with(test, IRanges(age_younger, age_older, names = species)) # 构建时间分箱的IRanges对象(0-5000 BP,每500年一个分箱) ir2 = IRanges(start = seq(0,4501,500), end = seq(500,5000,500)) # 计算分箱与样本的重叠数(非唯一物种数) overlaps <- ir2 %>% mutate(n_overlaps = count_overlaps(., ir1)) overlaps
输出结果:
IRanges object with 10 ranges and 1 metadata column: start end width | n_overlaps <integer> <integer> <integer> | <integer> [1] 0 500 501 | 3 [2] 500 1000 501 | 4 [3] 1000 1500 501 | 4 [4] 1500 2000 501 | 5 [5] 2000 2500 501 | 3 [6] 2500 3000 501 | 1 [7] 3000 3500 501 | 1 [8] 3500 4000 501 | 4 [9] 4000 4500 501 | 4 [10] 4500 5000 501 | 1
解决方案:统计分箱内唯一物种数
使用findOverlaps获取分箱与样本区间的重叠配对,结合dplyr分组去重计数,最终转换为目标数据框:
library(IRanges) library(dplyr) # 构建物种时间范围IRanges,保留物种信息作为元数据 ir1 <- IRanges( start = test$age_younger, end = test$age_older, species = test$species ) # 构建时间分箱IRanges,添加分箱标识(用起始值作为age_bin) ir2 <- IRanges( start = seq(0, 4501, 500), end = seq(500, 5000, 500), age_bin = seq(0, 4500, 500) ) # 找到所有分箱与样本区间的重叠配对 olaps <- findOverlaps(ir2, ir1) # 关联分箱、物种信息,统计每个分箱的唯一物种数 result <- as.data.frame(olaps) %>% left_join(as.data.frame(ir2) %>% select(queryHits = seqnames, age_bin), by = "queryHits") %>% left_join(as.data.frame(ir1) %>% select(subjectHits = seqnames, species), by = "subjectHits") %>% group_by(age_bin) %>% summarise(species_count = n_distinct(species)) %>% arrange(age_bin) # 查看结果 result
输出结果(针对示例数据集)
# A tibble: 10 × 2 age_bin species_count <dbl> <int> 1 0 2 2 500 4 3 1000 4 4 1500 4 5 2000 3 6 2500 1 7 3000 1 8 3500 3 9 4000 4 10 4500 1
代码说明
- 构建
ir1时直接将species作为元数据存入IRanges,避免依赖名称关联的潜在问题; - 给
ir2添加age_bin元数据,直接用分箱的起始值作为标识,符合输出需求; findOverlaps返回所有分箱与样本区间的配对关系,确保无遗漏;- 通过两次
left_join将分箱、物种信息关联,再用n_distinct统计每个分箱的唯一物种数; - 最终结果按
age_bin排序,适配时间线图表的展示逻辑; - 天然支持负年龄区间和零长度区间的重叠判断。
内容的提问来源于stack exchange,提问作者Mattias Sjölander
相关产品推荐
相关产品推荐

