You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R补全基因组覆盖度数据框缺失位点并填充0

补全基因组位点覆盖度数据(支持多基因组场景)

需求说明

现有包含基因组(genome)、位点(position)、覆盖度(coverage)、**基因组长度(length)**信息的数据框df,原始数据仅记录了覆盖度不为0的位点。需补全每个基因组的所有位点:从position=1到length列指定的数值,缺失位点的coverage填充为0,同时支持多基因组并行处理。

示例输入

单基因组输入

> df
    genome position coverage length
1  NC_2424        3        1     30
2  NC_2424        5        1     30
3  NC_2424        6        1     30
4  NC_2424        7        1     30
5  NC_2424        8        4     30
6  NC_2424       14        4     30
7  NC_2424       15        6     30
8  NC_2424       16        2     30
9  NC_2424       20        3     30
10 NC_2424       21        1     30

多基因组输入

> df
    genome position coverage length
1  NC_2424        3        1     30
2  NC_2424        5        1     30
3  NC_2424        6        1     30
4  NC_2424        7        1     30
5  NC_2424        8        4     30
6  NC_35131       14        4     34
7  NC_35131       15        6     34
8  NC_35131       16        2     34
9  NC_35131       20        3     34
10 NC_35131       21        1     34

解决方案代码

使用dplyr分组、tidyr::complete补全缺失位点:

df %>%
  dplyr::group_by(genome) %>%
  tidyr::complete(genome, position = seq(as.integer(unique(length))), length, fill = list(coverage = 0))

示例输出(单基因组补全结果)

> df.out
    genome position coverage length
1  NC_2424        1        0     30
2  NC_2424        2        0     30
3  NC_2424        3        1     30
4  NC_2424        4        0     30
5  NC_2424        5        1     30
6  NC_2424        6        1     30
7  NC_2424        7        1     30
8  NC_2424        8        4     30
9  NC_2424        9        0     30
10 NC_2424       10        0     30
11 NC_2424       11        0     30
12 NC_2424       12        0     30
13 NC_2424       13        0     30
14 NC_2424       14        4     30
15 NC_2424       15        6     30
16 NC_2424       16        2     30
17 NC_2424       17        0     30
18 NC_2424       18        0     30
19 NC_2424       19        0     30
20 NC_2424       20        3     30
21 NC_2424       21        1     30
22 NC_2424       22        0     30
23 NC_2424       23        0     30
24 NC_2424       24        0     30
25 NC_2424       25        0     30
26 NC_2424       26        0     30
27 NC_2424       27        0     30
28 NC_2424       28        0     30
29 NC_2424       29        0     30
30 NC_2424       30        0     30

可复现数据结构

原始数据结构

> dput(df)
structure(list(genome = c("NC_2424", "NC_2424", "NC_2424", "NC_2424", 
"NC_2424", "NC_2424", "NC_2424", "NC_2424", "NC_2424", "NC_2424"
), position = c(3, 5, 6, 7, 8, 14, 15, 16, 20, 21), coverage = c(1, 
1, 1, 1, 4, 4, 6, 2, 3, 1), length = c("30", "30", "30", "30", 
"30", "30", "30", "30", "30", "30")), class = "data.frame", row.names = c(NA, 
-10L))

补全后数据结构

> dput(df.out)
structure(list(genome = c("NC_2424", "NC_2424", "NC_2424", "NC_2424", 
"NC_2424", "NC_2424", "NC_2424", "NC_2424", "NC_2424", "NC_2424", 
"NC_2424", "NC_2424", "NC_2424", "NC_2424", "NC_2424", "NC_2424", 
"NC_2424", "NC_2424", "NC_2424", "NC_2424", "NC_2424", "NC_2424", 
"NC_2424", "NC_2424", "NC_2424", "NC_2424", "NC_2424", "NC_2424", 
"NC_2424", "NC_2424"), position = 1:30, coverage = c(0, 0, 1, 
0, 1, 1, 1, 4, 0, 0, 0, 0, 0, 4, 6, 2, 0, 0, 0, 3, 1, 0, 0, 0, 
0, 0, 0, 0, 0, 0), length = c("30", "30", "30", "30", "30", "30", 
"30", "30", "30", "30", "30", "30", "30", "30", "30", "30", "30", 
"30", "30", "30", "30", "30", "30", "30", "30", "30", "30", "30", 
"30", "30")), class = "data.frame", row.names = c(NA, -30L))

内容的提问来源于stack exchange,提问作者user2300940

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 19:20:40