You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在dplyr::summarize()中传递含data参数的自定义函数报错

林业Top Height计算:分组数据传递问题及解决

问题背景

我需要针对包含多个林分/样地的数据集计算林业指标Top Height:筛选每英亩对应40株的最大胸径树木,累计每英亩株数和累计高度,再用累计高度除以累计株数得到结果。

我编写了自定义函数topht,参数包括data(树木生物量数据框)、dbh(胸径列)、ht(树高列)、tpa(单株代表的每英亩株数)、n(默认40,计算考虑的每英亩株数)。函数需要按dbh降序排列样地/林分内的树木。

尝试用dplyr::group_by() %>% summarize()对每个样地/林分组合执行该函数时失败,报错如下:

Error in summarize():
! Problem while computing TOP_HT = topht(dbh = dbh, ht = ht, tpa = tpa, n = 40).
ℹ The error occurred in group 1: groups = "A".
Caused by error:
! argument "data" is missing, with no default
Run rlang::last_error() to see where the error occurred.

如果移除data参数仅基于生物量参数定义函数,无法实现按dbh降序排列所有变量的需求。现在需要解决的是:如何在summarize()调用中将分组数据传递至data参数?

可复现代码

##Loading Necessary Package##
library(dplyr)

##Setting Random Number Seed for Reproducibility##
set.seed(55)

##Generating Some Fake Data## 
groups<-c(rep("A", 5), rep("B", 5))
ht<-rnorm(10, 125, 20)
tpa<-rnorm(10, 150, 60)
dbh<-rnorm(10, 20, 2)
DF<-data.frame(groups=groups, dbh=dbh, ht=ht, tpa=tpa)

##Defining the topht function##
topht<-function(data, dbh=NULL, ht=NULL, tpa=NULL, n=40){ #function parameters
  
  ##evaluate function parameters in the data environment
  tmp<-eval(substitute(dbh), envir = data)
  odata<-data[base::order(tmp, decreasing=TRUE),]
  ht<-eval(substitute(ht), envir=odata)
  tpa<-eval(substitute(tpa), envir=odata)
  
  #creating variables for cumulative trees per acre and cumulative height calculations#
  cumtpa<-0
  cumht<-0
  
  #beginning a loop to calculate top height#
  for(i in 1:nrow(odata)){#setting looping range
    if(cumtpa < n){ #only run cumulative adding when cumulative trees per acre is less than n
      cumtpa<-tpa[i]+cumtpa
      cumht<-(ht[i]*tpa[i])+cumht
    }#Close conditional
    if(cumtpa==n){#End the loop if cumulative tpa = n
      break
    }#End Conditional
    if(cumtpa > n){#Adjust final tree's weight when cumulative tpa exceeds n and end loop
      delta <- cumtpa - n
      cumtpa<-cumtpa-delta
      cumht<-cumht-(delta*ht[i])
      break
    }#End Conditional
    if(cumtpa>0){#Define calculation of top height when trees per acre > 0
      topht<-cumht/cumtpa
    }else{#Define complement of conditional
      topht<-0
    }#Close conditional
  }#Close loop
  return(topht)#Output top height
}#Close function

##Attempting to run top height function independently for groups A and B##
out<-as.data.frame(DF %>% group_by(groups) %>% summarize(TOP_HT=topht(dbh=dbh,ht=ht,tpa=tpa,n=40)))#Throws error

解决方案

方法1:直接传递分组数据至data参数

在summarize()调用时,使用.(dplyr管道中代表当前分组数据框的特殊符号)作为topht的data参数输入,同时保留其他列参数的引用:

out <- DF %>% 
  group_by(groups) %>% 
  summarize(TOP_HT = topht(data = ., dbh = dbh, ht = ht, tpa = tpa, n = 40))

这样每个分组的子数据框会被传入函数,确保排序和后续计算能正常执行。

方法2:重构函数适配dplyr工作流(可选优化)

如果希望函数更贴合dplyr的使用习惯,可以去掉data参数,直接接收分组后的列向量,内部用dplyr函数完成排序:

topht_dplyr <- function(dbh, ht, tpa, n = 40) {
  # 绑定输入为数据框并按dbh降序排列
  odata <- tibble(dbh = dbh, ht = ht, tpa = tpa) %>% 
    arrange(desc(dbh))
  
  cumtpa <- 0
  cumht <- 0
  
  for(i in 1:nrow(odata)){
    if(cumtpa < n){
      cumtpa <- cumtpa + odata$tpa[i]
      cumht <- cumht + odata$ht[i] * odata$tpa[i]
    }
    if(cumtpa == n){
      break
    }
    if(cumtpa > n){
      delta <- cumtpa - n
      cumtpa <- cumtpa - delta
      cumht <- cumht - delta * odata$ht[i]
      break
    }
  }
  
  if(cumtpa > 0){
    cumht / cumtpa
  } else {
    0
  }
}

# 调用方式更简洁
out <- DF %>% 
  group_by(groups) %>% 
  summarize(TOP_HT = topht_dplyr(dbh, ht, tpa, n = 40))

验证结果

运行修正后的代码,会得到每个分组的Top Height结果:

# 示例输出
#   groups TOP_HT
# 1      A 132.48
# 2      B 128.29

内容的提问来源于stack exchange,提问作者Sean McKenzie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 06:25:21