You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用datarium包weightloss数据集做RMANOVA分组Shapiro-Wilk检验报错求解

分组执行Shapiro-Wilk检验报错问题分析

背景

使用datarium包中的weightloss数据集运行重复测量方差分析,数据集前6行的dput输出如下:

dput(head(weightloss))
structure(list(id = structure(1:6, .Label = c("1", "2", "3", 
"4", "5", "6", "7", "8", "9", "10", "11", "12"), class = "factor"), 
    diet = structure(c(1L, 1L, 1L, 1L, 1L, 1L), .Label = c("no", 
    "yes"), class = "factor"), exercises = structure(c(1L, 1L, 
    1L, 1L, 1L, 1L), .Label = c("no", "yes"), class = "factor"), 
    t1 = c(10.43, 11.59, 11.35, 11.12, 9.5, 9.5), t2 = c(13.21, 
    10.66, 11.12, 9.5, 9.73, 12.74), t3 = c(11.59, 13.21, 11.35, 
    11.12, 12.28, 10.43)), row.names = c(NA, -6L), class = c("tbl_df", 
"tbl", "data.frame"))

已完成的操作

数据重塑与箱线图绘制

代码如下:

# Create Data Frame for Dataset:
weight <- weightloss
weight

# Pivot Longer Data to Create Factors and Scores:
weight <- weight %>% 
  pivot_longer(names_to = 'trial', # creates factor (x)
               values_to = 'value', # creates value (y)
               cols = t1:t3) # finds which cols to factor

# Plot Means in Boxplot:
ggplot(weight,
       aes(x=trial,y=value))+
  geom_boxplot()+
  labs(title = "Trial Means") # As can be predicted, inc w/time

运行后得到正常箱线图:
Boxplot

离群值识别与正态性初检

代码如下:

# Identify Outliers (Should be None Given Boxplot):
outlier <- weight %>% 
  group_by(trial) %>% 
  identify_outliers(value)
outlier_frame <- data.frame(outlier) 
outlier_frame # none found :)

# Normality (Shapiro-Wilk and QQPlot):
model <- lm(value~trial,
            data = weight) # creates model
shapiro_test(residuals(model)) # measures Shapiro
ggqqplot(residuals(model))+
  labs(title = "QQ Plot of Residuals") # creates QQ

运行后得到正常残差QQ图:
QQPLOT

分面QQ图绘制

代码如下:

ggqqplot(weight, "value", ggtheme = theme_bw())+
  facet_wrap(~trial)+
labs(title = "QQPlot of Each Trial") #looks normal

输出符合正态性预期:
QQPLOT FACETED

报错情况与原因分析

第一个报错

尝试的代码:

shapiro_group <- weight %>%
  group_by(trial) %>%
  shapiro_test(value)

报错内容:

错误:mutate()列data存在问题。i data = map(.data$data, .f, ...)。x 必须按.data中存在的变量分组。

  • 列variable不存在。

原因:rstatix包的shapiro_test()在旧版本中不直接支持dplyr::group_by()返回的分组tibble结构,函数内部没有适配分组后的数据迭代逻辑。
正确写法可选以下任意一种:

# 写法1:直接指定分组变量公式
shapiro_group <- shapiro_test(weight, value ~ trial)

# 写法2:结合do()适配分组逻辑
shapiro_group <- weight %>%
  group_by(trial) %>%
  do(shapiro_test(., value))

# 写法3:基础包搭配tidyr实现
library(tidyr)
shapiro_group <- weight %>%
  group_by(trial) %>%
  summarise(shapiro_res = list(shapiro.test(value))) %>%
  unnest_wider(shapiro_res)

第二个报错

尝试的代码:

shapiro_test(weight, trial$value)

报错内容:

错误:无法选取不存在的列。x 列trial$value不存在。

原因:trial是weight数据框内的列,不是环境中独立存在的数据框对象,trial$value的写法会尝试从独立的trial对象中提取value列,自然无法找到对应列。如果要指定数据框内的列,直接传入列名即可,不需要加数据框前缀。
正确写法:

shapiro_test(weight, value)

内容的提问来源于stack exchange,提问作者Shawn Hemelstrand

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 06:57:03