使用group_by后调用shapiro.test报错:unused argument (ratio_log)求助
Fixing the Shapiro-Wilk Test Error with
group_by The error you're seeing comes down to two key issues in your code:
- You didn't save the
ratio_logcolumn to your dataframe after usingmutate(). Themutate()function returns a new dataframe, but you didn't assign it back todf—so the column doesn't exist when you try to reference it in your pipe chain. shapiro.test()isn't a dplyr verb, so you can't pass a column name directly to it after grouping. You need to use a dplyr function that applies the test to each group individually.
Here are a few corrected, working approaches:
Option 1: Use group_modify for tidy test results
This method applies shapiro.test() to each group and returns the results in a clean, structured dataframe:
library(dplyr) # Create dataframe and add calculated columns in one pipe df <- data.frame( Type = c("Bark", "Redwood", "Oak"), size = c(10,15,13), width = c(3,4,5) ) %>% mutate(Ratio = size/width, ratio_log = log10(Ratio)) # Run Shapiro-Wilk test per group df %>% group_by(Type) %>% group_modify(~ as.data.frame(shapiro.test(.$ratio_log)))
Option 2: Capture full test objects with summarize
If you want to keep the complete test objects for further inspection, use summarize() with a list, then unnest the results:
library(dplyr) library(tidyr) df %>% group_by(Type) %>% summarize(shapiro_result = list(shapiro.test(ratio_log))) %>% unnest_wider(shapiro_result)
Important Note
Your sample dataset only has one observation per group, which is too small for a Shapiro-Wilk test (it requires at least 3 observations to produce meaningful results). Make sure your actual dataset has enough data points per group before running this test.
内容的提问来源于stack exchange,提问作者Brandon Jablon
相关产品推荐
相关产品推荐

