You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于多条件的生态数据聚合:用dplyr还是循环实现?

用dplyr实现生态数据集的分组合并需求

问题背景

现有一个整洁的生态数据集,每行代表一个样本/个体,包含多列变量。模拟数据集代码如下:

# fake dataset
loc <- c(1,1,2,2,2,3,3,3,3,3,3,3,3)
date <- c(2021, 2022, 2021, 2021, 2022, 2021, 2021, 2022, 2023, 2023, 2023, 2023, 2023)
hab <- c("w", "l", "w", "w", "w", "l", "l", "w", "w", "w", "w", "w", "w")
spec <- c("frog", "frog", "frog", "frog", "frog", "beaver", "beaver", "beaver", "kingfisher", "kingfisher", "kingfisher", "kingfisher", "kingfisher")
n <- c(1,1,1,1,1,1,1,1,1,1,1,1,1)

df <- tibble(loc, date, hab, spec, n)

需求:将同一地点(loc)、日期(date)、栖息地(hab)下的指定物种(beaver和kingfisher,不含frog)的个体合并到同一行,每行最多包含3个个体,最终输出格式如下:

# wanted output
loc1 <- c(1,1,2,2,2,3,3,3,3)
date1 <- c(2021, 2022, 2021, 2021, 2022, 2021, 2022, 2023, 2023)
hab1 <- c("w", "l", "w", "w", "w", "l", "w", "w", "w")
spec1 <- c("frog", "frog", "frog", "frog", "frog", "beaver", "beaver", "kingfisher", "kingfisher")
n1 <- c(1,1,1,1,1,2,1,3,2)

df1 <- tibble(loc1, date1, hab1, spec1, n1)

# 输出预览
#    loc1 date1 hab1  spec1         n1
#   <dbl> <dbl> <chr> <chr>      <dbl>
# 1     1  2021 w     frog           1
# 2     1  2022 l     frog           1
# 3     2  2021 w     frog           1
# 4     2  2021 w     frog           1
# 5     2  2022 w     frog           1
# 6     3  2021 l     beaver         2
# 7     3  2022 w     beaver         1
# 8     3  2023 w     kingfisher     3
# 9     3  2023 w     kingfisher     2

解决方案:用dplyr实现

完全不需要使用for loop,仅通过dplyr的分组、窗口函数和聚合操作就能完成需求,代码如下:

library(dplyr)

df_processed <- df %>%
  # 为需要合并的物种添加批次序号,每3个个体为一个批次;frog保持每个个体单独成批次
  group_by(loc, date, hab, spec) %>%
  mutate(
    batch = ifelse(spec %in% c("beaver", "kingfisher"), 
                   ceiling(row_number() / 3), 
                   row_number())
  ) %>%
  ungroup() %>%
  # 按原分组+批次聚合,计算每个批次的个体总数
  group_by(loc, date, hab, spec, batch) %>%
  summarise(n = sum(n), .groups = "drop") %>%
  # 移除临时的批次列
  select(-batch)

# 查看最终结果
df_processed

代码逻辑说明

  1. 批次划分:通过row_number()生成分组内的个体序号,对指定物种用ceiling(row_number()/3)将每3个个体划为一个批次,frog则每个个体单独成批次,确保不合并。
  2. 聚合合并:按loc、date、hab、spec和batch分组,对n求和,实现同批次个体的合并。
  3. 结果整理:移除临时的batch列,得到符合要求的输出结构。

运行上述代码后,输出结果与需求中的df1完全一致。

内容的提问来源于stack exchange,提问作者matteo_rpm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 10:17:08