在dplyr mutate中自定义case_when函数报错:找不到对象'WK1'
问题解决:动态NFL数据集按周计算指定列平均值
问题场景
现有随NFL赛季推进每周新增列的数据集,尝试基于当前周数自定义average16函数,在dplyr::mutate中调用时始终返回! object 'WK1' not found错误。
示例数据(第二周)
library(dplyr) library(stringr) NumCorrect <- data.frame( TEAM = c(str_c("team_", 1:5)), WK1 = c(11,12,11,12,13), WK2 = c(7,7,9,10,7) )
报错代码
# 定义当前周数 curWEEK = 2 # 缺失数据来源的mutate调用 correct_d <- mutate(AVG_16 = average16(curWEEK)) # 自定义函数 average16 <- function(x) {case_when(x == 1 ~ WK1, x == 2 ~ round(mean((WK1:WK2), na.rm=TRUE),1), x == 3 ~ round(mean((WK1:WK3), na.rm=TRUE),1), x %in% 4:7 ~ round(mean((WK1:WK4), na.rm=TRUE),1), x %in% 8:11 ~ round(mean(c(WK1:WK4,WK8), na.rm=TRUE),1), x %in% 12:14 ~ round(mean(c(WK1:WK4,WK8,WK12), na.rm=TRUE),1), x == 15 ~ round(mean(c(WK1:WK4,WK8,WK12,WK15), na.rm=TRUE),1), x == 16 ~ round(mean(c(WK1:WK4,WK8,WK12,WK15:WK16), na.rm=TRUE),1), x == 17 ~ round(mean(c(WK1:WK4,WK8,WK12,WK15:WK17), na.rm=TRUE),1), x == 18 ~ round(mean(c(WK1:WK4,WK8,WK12,WK15:WK18), na.rm=TRUE),1) )}
错误信息
Error in `mutate()`: ℹ In argument: `AVG_16 = average16(curWEEK)`. ℹ In row 1. Caused by error in `case_when()`: ! Failed to evaluate the right-hand side of formula 1. Caused by error: ! object 'WK1' not found
错误原因
- mutate调用缺失数据上下文:原代码中
mutate未指定处理的数据框,函数无法定位WK1等列。 - 函数未绑定数据环境:自定义函数直接引用列名,但函数自身环境中无这些列的定义,需传入数据框或利用tidyeval语法识别数据内的列。
- 列引用逻辑错误:
WK1:WK2是生成数值序列,而非选择数据框中的列,应改用列选择或合并列向量的方式。
解决方案
方法1:修改函数接收数据参数
调整函数结构,明确传入数据框和当前周数,动态生成列名并计算行均值:
average16 <- function(data, x) { case_when( x == 1 ~ data$WK1, x == 2 ~ round(rowMeans(data[, paste0("WK", 1:2)], na.rm = TRUE), 1), x == 3 ~ round(rowMeans(data[, paste0("WK", 1:3)], na.rm = TRUE), 1), x %in% 4:7 ~ round(rowMeans(data[, paste0("WK", 1:4)], na.rm = TRUE), 1), x %in% 8:11 ~ round(rowMeans(data[, c(paste0("WK", 1:4), "WK8")], na.rm = TRUE), 1), x %in% 12:14 ~ round(rowMeans(data[, c(paste0("WK", 1:4), "WK8", "WK12")], na.rm = TRUE), 1), x == 15 ~ round(rowMeans(data[, c(paste0("WK", 1:4), "WK8", "WK12", "WK15")], na.rm = TRUE), 1), x == 16 ~ round(rowMeans(data[, c(paste0("WK", 1:4), "WK8", "WK12", paste0("WK", 15:16))], na.rm = TRUE), 1), x == 17 ~ round(rowMeans(data[, c(paste0("WK", 1:4), "WK8", "WK12", paste0("WK", 15:17))], na.rm = TRUE), 1), x == 18 ~ round(rowMeans(data[, c(paste0("WK", 1:4), "WK8", "WK12", paste0("WK", 15:18))], na.rm = TRUE), 1) ) } # 正确调用mutate,传递数据框上下文 correct_d <- NumCorrect %>% mutate(AVG_16 = average16(., curWEEK)) # 查看结果 correct_d
方法2:用tidyeval语法适配动态列
利用dplyr的选择函数,灵活匹配赛季新增的列,避免硬编码:
average16 <- function(x) { case_when( x == 1 ~ pull(WK1), x == 2 ~ round(rowMeans(select(starts_with("WK") & matches(paste0("WK", 1:2))), na.rm = TRUE), 1), x == 3 ~ round(rowMeans(select(starts_with("WK") & matches(paste0("WK", 1:3))), na.rm = TRUE), 1), x %in% 4:7 ~ round(rowMeans(select(starts_with("WK") & matches(paste0("WK", 1:4))), na.rm = TRUE), 1), x %in% 8:11 ~ round(rowMeans(select(starts_with("WK") & matches(paste0("WK", c(1:4,8)))), na.rm = TRUE), 1), x %in% 12:14 ~ round(rowMeans(select(starts_with("WK") & matches(paste0("WK", c(1:4,8,12)))), na.rm = TRUE), 1), x == 15 ~ round(rowMeans(select(starts_with("WK") & matches(paste0("WK", c(1:4,8,12,15)))), na.rm = TRUE), 1), x == 16 ~ round(rowMeans(select(starts_with("WK") & matches(paste0("WK", c(1:4,8,12,15:16)))), na.rm = TRUE), 1), x == 17 ~ round(rowMeans(select(starts_with("WK") & matches(paste0("WK", c(1:4,8,12,15:17)))), na.rm = TRUE), 1), x == 18 ~ round(rowMeans(select(starts_with("WK") & matches(paste0("WK", c(1:4,8,12,15:18)))), na.rm = TRUE), 1) ) } # 调用方式 correct_d <- NumCorrect %>% mutate(AVG_16 = average16(curWEEK))
关键注意点
- 必须通过管道
%>%给mutate传递数据框,确保函数能获取列的上下文。 - 使用
rowMeans而非mean,因为需要对每行的多列计算平均值,mean会计算整个向量的全局均值。 - 用
paste0("WK", 1:2)动态生成列名,适配赛季推进后新增的列,避免重复硬编码。
内容的提问来源于stack exchange,提问作者D_Bugli
相关产品推荐
相关产品推荐

