R语言scout函数报错排查:condition has length >1及NA警告
问题排查与修复:足球球员指标筛选scout函数报错处理
问题背景
我有一份涵盖五大联赛所有足球球员的数据集,正在构建scout函数,用来筛选出所选指标处于前85百分位的球员名单。测试调用scout(Total_Big_5_new,"Nutmegs")时出现报错:
the condition has length > 1
In addition: Warning message:
In percentile(database$metric) : NAs introduced by coercion
现有代码
scout函数
scout <- function(database, ...) { l <- list(...) l2 <- list() j <- 1 for(metric in l){ if(metric %in% colnames(database)){ l2[[j]] <- percentile(database[[metric]]) j <- j + 1 }else{ print(paste("The stat", metric, "is not recorded")) } } i <- 1 k <- 1 shortlist <- list() for (player in database){ compared <- select(database, unlist(l)) if (all(compared) > all(unlist(l2))){ shortlist[[i]] <- player i <- i + 1 } } return(shortlist) }
percentile函数
percentile <- function(metric, value = 0.85) { answer <- unname(quantile(metric, c(value))) return(as.numeric(paste(answer))) }
示例数据集与预期结果
构造随机数据集df:
df <- as_tibble(data.frame( Player = c(LETTERS[1:13]), Goals = c(sample(1:45, 13, replace=FALSE)), Assists = c(sample(1:31, 13, replace=FALSE)), Nutmegs = c(sample(1:28, 13, replace = FALSE)), Dribbles = c(sample(43:208, 13, replace = FALSE)) ))
数据集内容:
Player Goals Assists Nutmegs Dribbles <chr> <int> <int> <int> <int> 1 A 23 16 1 125 2 B 7 2 19 195 3 C 21 4 28 142 4 D 28 19 23 112 5 E 8 27 26 152 6 F 17 23 16 45 7 G 30 6 25 206 8 H 26 24 8 136 9 I 18 3 27 99 10 J 31 25 7 198 11 K 4 21 13 82 12 L 1 13 22 66 13 M 43 7 4 194
调用percentile(df$Goals, 0.65)返回25.4,预期调用scout(df,"Goals")应返回球员D、G、H、J、M。
报错原因分析
- 遍历逻辑错误:
for (player in database)会遍历数据集的列而非行,导致后续筛选逻辑完全偏离预期。 - 条件判断逻辑错误:
all(compared) > all(unlist(l2))会把整列数据转为单一逻辑值,无法实现逐行比较每个球员的指标是否超过对应百分位阈值。 - 冗余数值转换:
as.numeric(paste(answer))属于多余操作,quantile返回的数值可直接使用,该操作是NA警告的潜在来源。
修复后的代码
修复percentile函数
percentile <- function(metric, value = 0.85) { # 直接返回百分位结果,添加NA值处理 unname(quantile(metric, value, na.rm = TRUE)) }
修复scout函数
library(dplyr) scout <- function(database, ...) { metrics <- unlist(list(...)) # 校验指标有效性 valid_metrics <- metrics[metrics %in% colnames(database)] invalid_metrics <- setdiff(metrics, valid_metrics) if(length(invalid_metrics) > 0){ lapply(invalid_metrics, function(m) { message(paste("The stat", m, "is not recorded")) }) } if(length(valid_metrics) == 0) { message("No valid metrics provided") return(data.frame()) } # 计算每个指标的85百分位阈值 thresholds <- sapply(valid_metrics, function(m) { percentile(database[[m]]) }) # 筛选所有指标达标球员 shortlist <- database %>% filter(across(all_of(valid_metrics), ~ .x >= thresholds[cur_column()])) return(shortlist) }
验证结果
调用scout(df, "Goals"),会自动计算Goals的85百分位阈值,筛选出所有进球数≥该阈值的球员,符合预期结果。
内容的提问来源于stack exchange,提问作者med0504
相关产品推荐
相关产品推荐

