You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言mgcv包gam函数中正确传递自定义weights vector

解决mgcv::gam在自定义函数中无法识别权重向量的问题

问题根源

mgcv包的gam函数在解析weights参数时,默认会在传入的data数据框或全局环境中查找对应变量,而非直接使用自定义函数的局部参数。当你在check_mod里传入weights_vector时,这个变量是函数的局部对象,gam的求值机制找不到它,因此抛出object 'weights_vector' not found错误。

解决方案

下面提供两种无需依赖全局环境变量名的可行方法:


方法1:临时将权重向量加入数据框

把权重向量临时添加到传入的data中,用完后再移除,避免污染原数据:

check_mod <- function(formula, dist, data, weights_vector = NULL, validation_data = NULL) {
  # 保存原数据列名,用于后续清理
  original_cols <- colnames(data)
  temp_weight_col <- ".temp_weights"
  
  if (!is.null(weights_vector)) {
    # 临时添加权重列到数据框
    data[[temp_weight_col]] <- weights_vector
    # gam中引用临时列
    model <- gam(formula, family = dist, data = data, method = "ML", weights = .temp_weights)
    # 移除临时列,恢复原数据结构
    data <- data[, original_cols, drop = FALSE]
  } else {
    model <- gam(formula, family = dist, data = data, method = "ML")
  }
  
  prediction_data <- if (!is.null(validation_data)) validation_data else data
  predicted <- predict(model, newdata = prediction_data)
  
  rmse <- sqrt(mean((prediction_data$Chl - predicted)^2))
  
  model_summary <- summary(model)
  
  deviance_explained <- model_summary$dev.expl
  degrees_of_freedom <- sum(model_summary$edf)
  ML <- model_summary[["sp.criterion"]][["ML"]]
  
  return(list(
    aic = AIC(model), 
    rmse = rmse, 
    dev_expl = deviance_explained, 
    edf = degrees_of_freedom, 
    ML = ML,
    model = model
  ))
}

方法2:使用do.call传递参数

通过do.call构建参数列表并调用gam,绕过gam的表达式求值逻辑,直接传递权重向量:

check_mod <- function(formula, dist, data, weights_vector = NULL, validation_data = NULL) {
  # 构建gam的基础参数列表
  gam_args <- list(
    formula = formula,
    family = dist,
    data = data,
    method = "ML"
  )
  
  # 如果有权重,添加到参数列表
  if (!is.null(weights_vector)) {
    gam_args$weights <- weights_vector
  }
  
  # 用do.call调用gam
  model <- do.call(gam, gam_args)
  
  prediction_data <- if (!is.null(validation_data)) validation_data else data
  predicted <- predict(model, newdata = prediction_data)
  
  rmse <- sqrt(mean((prediction_data$Chl - predicted)^2))
  
  model_summary <- summary(model)
  
  deviance_explained <- model_summary$dev.expl
  degrees_of_freedom <- sum(model_summary$edf)
  ML <- model_summary[["sp.criterion"]][["ML"]]
  
  return(list(
    aic = AIC(model), 
    rmse = rmse, 
    dev_expl = deviance_explained, 
    edf = degrees_of_freedom, 
    ML = ML,
    model = model
  ))
}

验证测试

使用你提供的复现代码调用修改后的check_mod,两种方法都能正常运行,不会再出现权重向量找不到的错误:

# 测试方法1或方法2的函数
res1 <- check_mod(formula_example, distribution_example, data_example, weights_vector = weights1)
res2 <- check_mod(formula_example, distribution_example, data_example, weights_vector = data_example$weights_col)

print(res1)
print(res2)

内容的提问来源于stack exchange,提问作者Paco Herrera

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 14:40:14