R函数调用GLM传入权重参数时出现‘对象未找到’错误求助
glm调用中weight对象找不到的问题解决
问题场景
给R包函数传入权重参数时出现错误:调用函数时通过weights = c("kernel_wght")指定权重,函数内部通过以下代码创建weight数据框:
weight1 <- sprintf("dataarg$%s", weights) weight <- as.data.frame(eval(parse(text = weight1)))
但执行glm调用时:
result1 <- glm(f1, family="gaussian", weights=weight, data=dataarg)
抛出错误:
Error in (function (arg) : object 'weight' not found
明明能打印weight确认其存在,但glm就是无法识别,需要明确原因和解决办法。
复现代码
模拟多重插补数据
dat <- c(1, 1, 0, .5, 1, 3, 0, 1, 1, 4, 0, .5, 1, 5, 1, 1, 1, 2, 1, .5, 2, 7, 1, 1, 2, 3, 0, .5, 2, 2, 0, 1, 2, 4, 1, .5) dat <- data.frame(matrix(dat,ncol=4, byrow=T)) colnames(dat) <- c("id", "y", "tx", "wt") imp_lst <- lapply(1:2, function(s) dplyr::filter(dat, id == s)) for (i in 1:length(imp_lst)) { assign(paste0("imp", i), as.data.frame(imp_lst[[i]])) } df_lst <- list() for (i in 1:length(imp_lst)) { assign(paste0("imp", i), as.data.frame(imp_lst[[i]])) df_lst <- append(df_lst, list(get(paste0("imp", i)))) names(df_lst)[i] <- paste0("imp", i) }
出错的示例函数
my_ex <- function(datasets, y, treatment, weights=NULL, ...) { data <- names(datasets) for (i in 1:length(treatment)) { d1 <- sprintf("datasets$%s", data[i]) dataarg <- eval(parse(text=d1)) print(dataarg) if(!is.null(weights)) { weight1 <- sprintf("dataarg$%s", weights) weight <- as.data.frame(eval(parse(text = weight1))) print(weight) } else { dataarg$weight <- weight <- rep(1,nrow(dataarg)) } f1 <- sprintf("%s ~ %s ", y, treatment) print(f1) result1 <- glm(f1, family="gaussian", weights=weight, data=dataarg) print(summary(result1)) } }
触发错误的调用
testrun <- my_ex(df_lst, y = c("y","y"), treatment = c("tx","tx"), weights = c("wt","wt"))
错误原因
- 环境解析问题:当
glm接收字符串形式的公式时,会默认在data参数指定的dataarg数据框环境中查找所有变量,包括weights参数。但你定义的weight是在函数的局部环境里,不在dataarg的环境范围内,所以glm找不到它。 - 参数类型错误:
glm的weights参数要求传入数值向量,但你把权重转成了数据框,这也是潜在的问题点。
解决办法
方法1:将权重向量加入数据框
把权重作为新列添加到dataarg中,让glm能在数据框环境里找到它:
my_ex <- function(datasets, y, treatment, weights=NULL, ...) { data <- names(datasets) for (i in 1:length(treatment)) { # 替换eval(parse),直接用列表索引获取数据 dataarg <- datasets[[data[i]]] print(dataarg) if(!is.null(weights)) { # 直接提取数值向量,不要转成数据框 weight <- dataarg[[weights[i]]] # 添加到数据框 dataarg$model_weight <- weight print(weight) } else { weight <- rep(1,nrow(dataarg)) dataarg$model_weight <- weight } f1 <- sprintf("%s ~ %s ", y[i], treatment[i]) print(f1) # 引用数据框中的权重列 result1 <- glm(f1, family="gaussian", weights=model_weight, data=dataarg) print(summary(result1)) } }
方法2:将字符串公式转为formula对象
formula对象会让glm优先在当前函数的局部环境中查找变量,同时把数据框转成向量传入:
my_ex <- function(datasets, y, treatment, weights=NULL, ...) { data <- names(datasets) for (i in 1:length(treatment)) { dataarg <- datasets[[data[i]]] print(dataarg) if(!is.null(weights)) { weight <- dataarg[[weights[i]]] print(weight) } else { weight <- rep(1,nrow(dataarg)) } # 转为formula对象 f1 <- as.formula(sprintf("%s ~ %s ", y[i], treatment[i])) print(f1) # 传入数值向量 result1 <- glm(f1, family="gaussian", weights=weight, data=dataarg) print(summary(result1)) } }
方法3:彻底移除eval(parse)
原代码中大量使用eval(parse)不仅容易引发环境问题,还降低了代码可读性和安全性,直接用列表/数据框索引替换即可,如上面两种方法中的dataarg <- datasets[[data[i]]]和weight <- dataarg[[weights[i]]]。
内容的提问来源于stack exchange,提问作者Jason Schoeneberger
相关产品推荐
相关产品推荐

