You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为分类层级的所有组/超集生成data.table行?

分类学数据集层级补全与求和问题

原始数据集

sample_table <- data.table(Phylum = c("Arthropoda", "Arthropoda", "Arthropoda", "Arthropoda"),
                          Class = c("Arachnida",   "Insecta",    "Insecta", "Insecta"),
                          Order = c("Acariformes",   "Coleoptera", "Coleoptera", "Coleoptera"),
                          Family = c(NA, "Staphylinidae", "Staphylinidae", "Staphylinidae"),
                          Genus = c(NA, "Staphylininae", "Staphylininae", "Staphylininae"),
                          Species = c(NA, NA, "Philonthus", "Philonthus"),
                          Site = c(5, 6, 6, 6),
                          Distance = c(0, 0, 0, 5),
                          N = c(58, 1, 5, 3))

# 数据预览
# Phylum       Class      Order        Family           Genus           Species     Site     Distance     N
# Arthropoda   Arachnida  Acariformes  <NA>             <NA>            <NA>        5        0            58
# Arthropoda   Insecta    Coleoptera   Staphylinidae    Staphylininae   <NA>        6        0            1
# Arthropoda   Insecta    Coleoptera   Staphylinidae    Staphylininae   Philonthus  6        0            5
# Arthropoda   Insecta    Coleoptera   Staphylinidae    Staphylininae   Philonthus  6        5            3

数据说明

Phylum、Class、Order、Family、Genus和Species为层级分类组,Site和Distance代表样本采集地点信息,N是对应分类组、站点及距离下的样本数量。

当前数据存在的问题:第2行(R2)的N值应包含第3行(R3)的N值,因为R2是R3的分类超集,但目前R2仅统计了无法鉴定到Species层级(即Species为NA)的样本。

期望输出数据集

# 调整后的数据预览
# Phylum       Class      Order        Family          Genus           Species     Site    Distance  N
# Arthropoda   <NA>       <NA>         <NA>            <NA>            <NA>        5       0         58
# Arthropoda   Arachnida  <NA>         <NA>            <NA>            <NA>        5       0         58
# Arthropoda   Arachnida  Acariformes  <NA>            <NA>            <NA>        5       0         58
# Arthropoda   <NA>       <NA>         <NA>            <NA>            <NA>        6       0         6
# Arthropoda   Insecta    <NA>         <NA>            <NA>            <NA>        6       0         6
# Arthropoda   Insecta    Coleoptera   <NA>            <NA>            <NA>        6       0         6
# Arthropoda   Insecta    Coleoptera   Staphylinidae   <NA>            <NA>        6       0         1
# Arthropoda   Insecta    Coleoptera   Staphylinidae   Staphylininae   <NA>        6       0         1
# Arthropoda   Insecta    Coleoptera   Staphylinidae   Staphylininae   Philonthus  6       0         5
# Arthropoda   <NA>       <NA>         <NA>            <NA>            <NA>        6       5         3
# Arthropoda   Insecta    <NA>         <NA>            <NA>            <NA>        6       5         3
# Arthropoda   Insecta    Coleoptera   <NA>            <NA>            <NA>        6       5         3
# Arthropoda   Insecta    Coleoptera   Staphylinidae   <NA>            <NA>        6       5         3
# Arthropoda   Insecta    Coleoptera   Staphylinidae   Staphylininae   <NA>        6       5         3
# Arthropoda   Insecta    Coleoptera   Staphylinidae   Staphylininae   Philonthus  6       5         3

调整后的数据特点:每个站点-距离对都对应所有分类超集的行,每行的N值包含鉴定到该行最高分类精度及更细精度的样本总量。

尝试过的方法与困惑

试过使用tidyr::complete()和交叉连接,但不确定是否适用。和常见示例不同的是,我希望在不跨不同分类分支的前提下补全数据集——本质是添加右侧分类列填充NA的行,但肯定有比手动操作更好的方法;同时也不清楚如何通过这些方法解决分组求和的问题。

编辑说明:样本数据集中新增了一行,用于说明Distance的处理方式(与Site一致,不同距离的样本不聚合)。

内容的提问来源于stack exchange,提问作者fisherpeak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 05:35:55