You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用dplyr::across按组执行mutate完成生物量标准化?

用dplyr实现按采样类型标准化生物量的方案

需求说明

原始生物数据每行代表一个个体,包含字段:

  • length_mm:体长数据
  • site:站点标识,前缀D代表潜水采样,G代表抓斗采样
  • log.afdw1至nonlin.sodw1:多列不同方法的生物量估计值

需要完成:

  1. 按站点汇总所有生物量字段的总和
  2. 根据采样类型(D/G)分别除以对应采样器面积(潜水用ek_area,抓斗用frame_area)
  3. 避免拆分数据后再合并的繁琐操作,全程用dplyr管道实现

实现代码(两种场景)

场景1:采样器面积存储在单独的映射表中

假设你有一个sampler_areas表,记录采样类型对应的面积:

# 示例采样器面积表
sampler_areas <- tibble(
  sampler_type = c("D", "G"),
  area = c(0.25, 0.1) # ek_area=0.25,frame_area=0.1
)

处理代码:

library(dplyr)

standardized_biomass <- bio_data %>%
  # 从site字段提取采样类型前缀
  mutate(sampler_type = substr(site, 1, 1)) %>%
  # 按站点+采样类型分组,批量汇总所有生物量字段
  group_by(site, sampler_type) %>%
  summarise(across(log.afdw1:nonlin.sodw1, sum, na.rm = TRUE), .groups = "drop") %>%
  # 关联采样器面积数据
  left_join(sampler_areas, by = "sampler_type") %>%
  # 批量标准化生物量,生成带_standardized后缀的新字段
  mutate(across(log.afdw1:nonlin.sodw1, ~ .x / area, .names = "{.col}_standardized")) %>%
  # 可选:保留核心结果字段
  select(site, sampler_type, ends_with("_standardized"))

场景2:采样器面积是固定常量

如果ek_area和frame_area是已知固定值,直接用case_when匹配:

library(dplyr)

# 定义固定采样器面积
ek_area <- 0.25
frame_area <- 0.1

standardized_biomass <- bio_data %>%
  mutate(sampler_type = substr(site, 1, 1)) %>%
  group_by(site, sampler_type) %>%
  summarise(across(log.afdw1:nonlin.sodw1, sum, na.rm = TRUE), .groups = "drop") %>%
  mutate(
    across(log.afdw1:nonlin.sodw1, 
           ~ case_when(
             sampler_type == "D" ~ .x / ek_area,
             sampler_type == "G" ~ .x / frame_area
           ),
           .names = "{.col}_standardized")
  )

关键逻辑说明

  • substr(site, 1, 1):直接从站点标识提取采样类型,无需拆分数据集
  • across():批量处理多列生物量字段,避免重复编写求和/标准化代码
  • 全程用dplyr管道串联操作,保持代码简洁且易维护,不需要拆分再合并数据

内容的提问来源于stack exchange,提问作者dandrews

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 05:01:01