如何使用dplyr::across按组执行mutate完成生物量标准化?
用dplyr实现按采样类型标准化生物量的方案
需求说明
原始生物数据每行代表一个个体,包含字段:
length_mm:体长数据site:站点标识,前缀D代表潜水采样,G代表抓斗采样log.afdw1至nonlin.sodw1:多列不同方法的生物量估计值
需要完成:
- 按站点汇总所有生物量字段的总和
- 根据采样类型(D/G)分别除以对应采样器面积(潜水用
ek_area,抓斗用frame_area) - 避免拆分数据后再合并的繁琐操作,全程用dplyr管道实现
实现代码(两种场景)
场景1:采样器面积存储在单独的映射表中
假设你有一个sampler_areas表,记录采样类型对应的面积:
# 示例采样器面积表 sampler_areas <- tibble( sampler_type = c("D", "G"), area = c(0.25, 0.1) # ek_area=0.25,frame_area=0.1 )
处理代码:
library(dplyr) standardized_biomass <- bio_data %>% # 从site字段提取采样类型前缀 mutate(sampler_type = substr(site, 1, 1)) %>% # 按站点+采样类型分组,批量汇总所有生物量字段 group_by(site, sampler_type) %>% summarise(across(log.afdw1:nonlin.sodw1, sum, na.rm = TRUE), .groups = "drop") %>% # 关联采样器面积数据 left_join(sampler_areas, by = "sampler_type") %>% # 批量标准化生物量,生成带_standardized后缀的新字段 mutate(across(log.afdw1:nonlin.sodw1, ~ .x / area, .names = "{.col}_standardized")) %>% # 可选:保留核心结果字段 select(site, sampler_type, ends_with("_standardized"))
场景2:采样器面积是固定常量
如果ek_area和frame_area是已知固定值,直接用case_when匹配:
library(dplyr) # 定义固定采样器面积 ek_area <- 0.25 frame_area <- 0.1 standardized_biomass <- bio_data %>% mutate(sampler_type = substr(site, 1, 1)) %>% group_by(site, sampler_type) %>% summarise(across(log.afdw1:nonlin.sodw1, sum, na.rm = TRUE), .groups = "drop") %>% mutate( across(log.afdw1:nonlin.sodw1, ~ case_when( sampler_type == "D" ~ .x / ek_area, sampler_type == "G" ~ .x / frame_area ), .names = "{.col}_standardized") )
关键逻辑说明
substr(site, 1, 1):直接从站点标识提取采样类型,无需拆分数据集across():批量处理多列生物量字段,避免重复编写求和/标准化代码- 全程用dplyr管道串联操作,保持代码简洁且易维护,不需要拆分再合并数据
内容的提问来源于stack exchange,提问作者dandrews
相关产品推荐
相关产品推荐

