使用datarium包weightloss数据集做RMANOVA分组Shapiro-Wilk检验报错求解
分组执行Shapiro-Wilk检验报错问题分析
背景
使用datarium包中的weightloss数据集运行重复测量方差分析,数据集前6行的dput输出如下:
dput(head(weightloss)) structure(list(id = structure(1:6, .Label = c("1", "2", "3", "4", "5", "6", "7", "8", "9", "10", "11", "12"), class = "factor"), diet = structure(c(1L, 1L, 1L, 1L, 1L, 1L), .Label = c("no", "yes"), class = "factor"), exercises = structure(c(1L, 1L, 1L, 1L, 1L, 1L), .Label = c("no", "yes"), class = "factor"), t1 = c(10.43, 11.59, 11.35, 11.12, 9.5, 9.5), t2 = c(13.21, 10.66, 11.12, 9.5, 9.73, 12.74), t3 = c(11.59, 13.21, 11.35, 11.12, 12.28, 10.43)), row.names = c(NA, -6L), class = c("tbl_df", "tbl", "data.frame"))
已完成的操作
数据重塑与箱线图绘制
代码如下:
# Create Data Frame for Dataset: weight <- weightloss weight # Pivot Longer Data to Create Factors and Scores: weight <- weight %>% pivot_longer(names_to = 'trial', # creates factor (x) values_to = 'value', # creates value (y) cols = t1:t3) # finds which cols to factor # Plot Means in Boxplot: ggplot(weight, aes(x=trial,y=value))+ geom_boxplot()+ labs(title = "Trial Means") # As can be predicted, inc w/time
运行后得到正常箱线图:
离群值识别与正态性初检
代码如下:
# Identify Outliers (Should be None Given Boxplot): outlier <- weight %>% group_by(trial) %>% identify_outliers(value) outlier_frame <- data.frame(outlier) outlier_frame # none found :) # Normality (Shapiro-Wilk and QQPlot): model <- lm(value~trial, data = weight) # creates model shapiro_test(residuals(model)) # measures Shapiro ggqqplot(residuals(model))+ labs(title = "QQ Plot of Residuals") # creates QQ
运行后得到正常残差QQ图:
分面QQ图绘制
代码如下:
ggqqplot(weight, "value", ggtheme = theme_bw())+ facet_wrap(~trial)+ labs(title = "QQPlot of Each Trial") #looks normal
输出符合正态性预期:
报错情况与原因分析
第一个报错
尝试的代码:
shapiro_group <- weight %>% group_by(trial) %>% shapiro_test(value)
报错内容:
错误:
mutate()列data存在问题。idata = map(.data$data, .f, ...)。x 必须按.data中存在的变量分组。
- 列
variable不存在。
原因:rstatix包的shapiro_test()在旧版本中不直接支持dplyr::group_by()返回的分组tibble结构,函数内部没有适配分组后的数据迭代逻辑。
正确写法可选以下任意一种:
# 写法1:直接指定分组变量公式 shapiro_group <- shapiro_test(weight, value ~ trial) # 写法2:结合do()适配分组逻辑 shapiro_group <- weight %>% group_by(trial) %>% do(shapiro_test(., value)) # 写法3:基础包搭配tidyr实现 library(tidyr) shapiro_group <- weight %>% group_by(trial) %>% summarise(shapiro_res = list(shapiro.test(value))) %>% unnest_wider(shapiro_res)
第二个报错
尝试的代码:
shapiro_test(weight, trial$value)
报错内容:
错误:无法选取不存在的列。x 列
trial$value不存在。
原因:trial是weight数据框内的列,不是环境中独立存在的数据框对象,trial$value的写法会尝试从独立的trial对象中提取value列,自然无法找到对应列。如果要指定数据框内的列,直接传入列名即可,不需要加数据框前缀。
正确写法:
shapiro_test(weight, value)
内容的提问来源于stack exchange,提问作者Shawn Hemelstrand
相关产品推荐
相关产品推荐

