如何为分类层级的所有组/超集生成data.table行?
分类学数据集层级补全与求和问题
原始数据集
sample_table <- data.table(Phylum = c("Arthropoda", "Arthropoda", "Arthropoda", "Arthropoda"), Class = c("Arachnida", "Insecta", "Insecta", "Insecta"), Order = c("Acariformes", "Coleoptera", "Coleoptera", "Coleoptera"), Family = c(NA, "Staphylinidae", "Staphylinidae", "Staphylinidae"), Genus = c(NA, "Staphylininae", "Staphylininae", "Staphylininae"), Species = c(NA, NA, "Philonthus", "Philonthus"), Site = c(5, 6, 6, 6), Distance = c(0, 0, 0, 5), N = c(58, 1, 5, 3)) # 数据预览 # Phylum Class Order Family Genus Species Site Distance N # Arthropoda Arachnida Acariformes <NA> <NA> <NA> 5 0 58 # Arthropoda Insecta Coleoptera Staphylinidae Staphylininae <NA> 6 0 1 # Arthropoda Insecta Coleoptera Staphylinidae Staphylininae Philonthus 6 0 5 # Arthropoda Insecta Coleoptera Staphylinidae Staphylininae Philonthus 6 5 3
数据说明
Phylum、Class、Order、Family、Genus和Species为层级分类组,Site和Distance代表样本采集地点信息,N是对应分类组、站点及距离下的样本数量。
当前数据存在的问题:第2行(R2)的N值应包含第3行(R3)的N值,因为R2是R3的分类超集,但目前R2仅统计了无法鉴定到Species层级(即Species为NA)的样本。
期望输出数据集
# 调整后的数据预览 # Phylum Class Order Family Genus Species Site Distance N # Arthropoda <NA> <NA> <NA> <NA> <NA> 5 0 58 # Arthropoda Arachnida <NA> <NA> <NA> <NA> 5 0 58 # Arthropoda Arachnida Acariformes <NA> <NA> <NA> 5 0 58 # Arthropoda <NA> <NA> <NA> <NA> <NA> 6 0 6 # Arthropoda Insecta <NA> <NA> <NA> <NA> 6 0 6 # Arthropoda Insecta Coleoptera <NA> <NA> <NA> 6 0 6 # Arthropoda Insecta Coleoptera Staphylinidae <NA> <NA> 6 0 1 # Arthropoda Insecta Coleoptera Staphylinidae Staphylininae <NA> 6 0 1 # Arthropoda Insecta Coleoptera Staphylinidae Staphylininae Philonthus 6 0 5 # Arthropoda <NA> <NA> <NA> <NA> <NA> 6 5 3 # Arthropoda Insecta <NA> <NA> <NA> <NA> 6 5 3 # Arthropoda Insecta Coleoptera <NA> <NA> <NA> 6 5 3 # Arthropoda Insecta Coleoptera Staphylinidae <NA> <NA> 6 5 3 # Arthropoda Insecta Coleoptera Staphylinidae Staphylininae <NA> 6 5 3 # Arthropoda Insecta Coleoptera Staphylinidae Staphylininae Philonthus 6 5 3
调整后的数据特点:每个站点-距离对都对应所有分类超集的行,每行的N值包含鉴定到该行最高分类精度及更细精度的样本总量。
尝试过的方法与困惑
试过使用tidyr::complete()和交叉连接,但不确定是否适用。和常见示例不同的是,我希望在不跨不同分类分支的前提下补全数据集——本质是添加右侧分类列填充NA的行,但肯定有比手动操作更好的方法;同时也不清楚如何通过这些方法解决分组求和的问题。
编辑说明:样本数据集中新增了一行,用于说明Distance的处理方式(与Site一致,不同距离的样本不聚合)。
内容的提问来源于stack exchange,提问作者fisherpeak
相关产品推荐
相关产品推荐

