在dplyr的across中传递自定义函数参数时遇错误求助
解决dplyr across中自定义函数传参的"object not found"错误
问题原因
当把lambda表达式替换为自定义函数get_which_max后,函数内的get(col_max)默认在函数自身的环境中查找对象,而非dplyr处理的分组数据框环境。而原lambda表达式直接在dplyr的tidy eval上下文执行,能正确识别数据框中的列,因此不会报错。
两种可行解决方案
方案1:使用.data代词明确指定数据框环境
修改自定义函数,用.data[[col_max]]替代get(col_max),强制从当前分组的数据框中提取参考列:
data(iris) ref_col <- "Sepal.Length" get_which_max <- function(x, col_max) { x[which.max(.data[[col_max]])] } iris_summary <- iris %>% group_by(Species) %>% summarise( Sepal.Length_max = max(Sepal.Length), across( Sepal.Width:Petal.Width, ~ get_which_max(.x, ref_col) ) )
方案2:传递整个分组数据框给自定义函数
让自定义函数接收完整的分组数据框,直接提取目标列和参考列进行计算,避免环境查找问题:
data(iris) ref_col <- "Sepal.Length" get_which_max <- function(df, target_col, ref_col) { df[[target_col]][which.max(df[[ref_col]])] } iris_summary <- iris %>% group_by(Species) %>% summarise( Sepal.Length_max = max(Sepal.Length), across( Sepal.Width:Petal.Width, ~ get_which_max(cur_data(), .y, ref_col), .names = "{col}_at_max_ref" ) )
注:cur_data()用于获取当前分组的数据框,.y代表当前遍历的列名,需要dplyr 1.0.0及以上版本支持。
内容的提问来源于stack exchange,提问作者emaoca
相关产品推荐
相关产品推荐

