按类别统计缺失值:R语言tidyverse代码报错原因咨询
is.na() Error in Your Tidyverse Code Hey there! Let's break down why your code is throwing that error and get it working correctly.
The Root Cause of the Error
The error message says "0 arguments passed to 'is.na' which requires 1" — that's exactly the problem! The is.na() function needs to know which column(s) you want to check for missing values, but right now you're calling it without passing any variable to it. R has no way to guess which data you want to evaluate, so it throws an error.
Corrected Code Options
Option 1: Total Missing Values Per Class (All Variables Combined)
If you want to calculate the total number of missing values across all variables for each Class, use across() to target all non-grouped columns (in this case, var1 and var2):
library(tidyverse) fakeData <- data.frame( var1 = c(1,2,NA,4,NA,6,7,8,9,10), var2=c(11,NA,NA,14,NA,16,17,NA,19,NA), Class = c(rep("A", 5), rep("B", 5)) ) fakeData %>% group_by(Class) %>% summarize(numMissing = sum(is.na(across(everything()))))
This will return:
# A tibble: 2 × 2 Class numMissing <chr> <int> 1 A 5 2 B 2
Option 2: Missing Values Per Variable, Per Class
If you want to see the number of missing values for each individual variable within each Class, adjust the code to return separate counts for var1 and var2:
fakeData %>% group_by(Class) %>% summarize( across( c(var1, var2), ~sum(is.na(.x)), .names = "numMissing_{.col}" ) )
This will give you:
# A tibble: 2 × 3 Class numMissing_var1 numMissing_var2 <chr> <int> <int> 1 A 2 3 2 B 0 2
How This Works
across(everything())tells dplyr to apply theis.na()check to every column except the grouped column (Class).~sum(is.na(.x))is a shorthand function where.xrepresents the current column being processed byacross().- The
.namesargument in Option 2 lets you customize the output column names for clarity.
内容的提问来源于stack exchange,提问作者buzaku

