在R中使用across对数据框多列应用自定义函数报错求助
问题描述
我有一个名为ex_ds的数据框,包含变量v1至v15:
> head(ex_ds) v1 v2 v3 v4 v5 v6 v7 v8 v9 v10 v11 v12 v13 v14 v15 1 5 2014 1 2 4 1 1 8 4 2 2 2 2 2 2 2 5 2014 2 6 1 3 8 <NA> 1 1 2 2 2 1 2 3 5 2014 2 5 2 1 1 8 1 1 1 <NA> <NA> 1 2 4 5 2014 2 2 1 4 1 2 5 2 2 2 2 2 2 5 5 2014 1 5 2 1 8 7 4 1 1 2 2 2 2 6 5 2014 2 4 3 5 3 1 3 2 2 2 2 2 2
我需要按v1和v2的组合分组,统计v3至v15每个变量的非NA响应数量。
我写了自定义函数,单独调用时可以正常运行:
n_fn_ex = function(var){ n_var = ex_ds %>% drop_na(var) %>% group_by(v1,v2) %>% count() }
单独调用结果:
> n_fn_ex("v3") # A tibble: 1 x 3 # Groups: v1, v2 [1] v1 v2 n <int> <int> <int> 1 5 2014 6 > n_fn_ex("v8") # A tibble: 1 x 3 # Groups: v1, v2 [1] v1 v2 n <int> <int> <int> 1 5 2014 5
但将其用于across语句时却报错:
ex_ds_1 = ex_ds %>% reframe(across(v3:v15, n_fn_ex))
错误信息:
Error in `reframe()`: i In argument: `across(v3:v15, n_fn_ex)`. Caused by error in `across()`: ! Can't compute column `v3`. Caused by error in `drop_na()`: ! Can't subset columns that don't exist. x Columns `1`, `2`, `2`, `2`, `1`, etc. don't exist. Run `rlang::last_trace()` to see where the error occurred.
我尝试过指定字符串列名、修改函数返回值等写法,均出现相同错误。
示例数据集:
ex_ds = structure(list(v1 = c(5L, 5L, 5L, 5L, 5L, 5L), v2 = c(2014L, 2014L, 2014L, 2014L, 2014L, 2014L), v3 = structure(c(1L, 2L, 2L, 2L, 1L, 2L), .Label = c("1", "2"), class = "factor"), v4 = structure(c(2L, 6L, 5L, 2L, 5L, 4L), .Label = c("1", "2", "3", "4", "5", "6"), class = "factor"), v5 = structure(c(4L, 1L, 2L, 1L, 2L, 3L), .Label = c("1", "2", "3", "4"), class = "factor"), v6 = structure(c(1L, 3L, 1L, 4L, 1L, 5L), .Label = c("1", "2", "3", "4", "5", "6"), class = "factor"), v7 = structure(c(1L, 8L, 1L, 1L, 8L, 3L), .Label = c("1", "2", "3", "4", "5", "6", "7", "8"), class = "factor"), v8 = structure(c(8L, NA, 8L, 2L, 7L, 1L), .Label = c("1", "2", "3", "4", "5", "6", "7", "8"), class = "factor"), v9 = structure(c(4L, 1L, 1L, 5L, 4L, 3L), .Label = c("1", "2", "3", "4", "5"), class = "factor"), v10 = structure(c(2L, 1L, 1L, 2L, 1L, 2L), .Label = c("1", "2"), class = "factor"), v11 = structure(c(2L, 2L, 1L, 2L, 1L, 2L), .Label = c("1", "2"), class = "factor"), v12 = structure(c(2L, 2L, NA, 2L, 2L, 2L), .Label = c("1", "2"), class = "factor"), v13 = structure(c(2L, 2L, NA, 2L, 2L, 2L), .Label = c("1", "2"), class = "factor"), v14 = structure(c(2L, 1L, 1L, 2L, 2L, 2L), .Label = c("1", "2"), class = "factor"), v15 = structure(c(2L, 2L, 2L, 2L, 2L, 2L), .Label = c("1", "2"), class = "factor")), row.names = c(NA, 6L), class = "data.frame")
错误原因
across在迭代列时,传递给函数的是列的数值向量,而非列名字符串。你的自定义函数n_fn_ex期望接收列名字符串,但across传入的是列的实际值(比如v3的向量c(1,2,2,2,1,2)),导致drop_na(var)尝试删除这些值对应的列,自然找不到,引发报错。另外,函数直接引用全局环境的ex_ds,没有利用across传递的分组上下文,也是问题之一。
解决方案
方法1:直接在across中统计非NA数量
无需自定义函数,直接用sum(!is.na(.x))实现统计,简洁高效:
library(dplyr) ex_ds_1 <- ex_ds %>% group_by(v1, v2) %>% reframe(across(v3:v15, ~sum(!is.na(.x)), .names = "n_{.col}"))
方法2:修改自定义函数适配across
让函数接收列向量,直接统计非NA数量:
n_fn_ex_fix <- function(col_vec) { sum(!is.na(col_vec)) } ex_ds_1 <- ex_ds %>% group_by(v1, v2) %>% reframe(across(v3:v15, n_fn_ex_fix, .names = "n_{.col}"))
方法3:用summarise替代reframe
效果与上述方法一致,适合习惯用summarise的场景:
ex_ds_1 <- ex_ds %>% group_by(v1, v2) %>% summarise(across(v3:v15, ~sum(!is.na(.x))), .groups = "drop")
输出结果
三种方法都会得到按v1、v2分组,每个变量的非NA计数:
> ex_ds_1 # A tibble: 1 x 15 v1 v2 n_v3 n_v4 n_v5 n_v6 n_v7 n_v8 n_v9 n_v10 n_v11 n_v12 n_v13 n_v14 n_v15 <int> <int> <int> <int> <int> <int> <int> <int> <int> <int> <int> <int> <int> <int> <int> 1 5 2014 6 6 6 6 6 5 6 6 6 5 5 6 6
内容的提问来源于stack exchange,提问作者abrar
相关产品推荐
相关产品推荐

