You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决rstatix包中anova_test函数的重复测量ANOVA报错问题

重复测量ANOVA报错原因分析与解决

问题重现

执行以下重复测量ANOVA代码:

res.aov = anova_test(data_ex,
                     dv = D, 
                     wid = A, 
                     within = B,
                     between = C,
                     type = 3)

出现如下错误:

Run `rlang::last_trace()` to see where the error occurred.
> rlang::last_trace()
<error/rlang_error>
Error in `spread()`:
! Each row of output must be identified by a unique combination of keys.
ℹ Keys are shared for 402 rows
• 1, 2
• 19, 20
...(省略重复行号)
---
Backtrace:
     ▆
  1. ├─rstatix::anova_test(...)
  2. │ ├─... %>% .append_anova_class()
  3. │ └─rstatix (local) .anova_test(...)
  4. │   └─rstatix:::car_anova(.args)
  5. │     └─rstatix::factorial_design(...)
  6. │       └─... %>% as_tibble()
  7. ├─rstatix (local) .append_anova_class(.)
  8. ├─tibble::as_tibble(.)
  9. ├─tidyr::spread(., key = ".group.", value = dv)
 10. └─tidyr:::spread.data.frame(., key = ".group.", value = dv)
Run rlang::last_trace(drop = FALSE) to see 2 hidden frames.

可复现的示例数据:

structure(list(A = c("1", "1", "1", "1", "1", "1", "2", "2", 
"2", "2", "2", "2", "3", "3", "3", "3", "3", "3", "4", "4", "4", 
"4", "4", "4", "5", "5", "5", "5", "5", "5", "6", "6", "6", "6", 
"6", "6", "7", "7", "7", "7", "7", "7", "8", "8", "8", "8", "8", 
"8", "9", "9"), B = structure(c(1L, 2L, 3L, 1L, 2L, 3L, 1L, 2L, 
3L, 1L, 2L, 3L, 1L, 2L, 3L, 1L, 2L, 3L, 1L, 2L, 3L, 1L, 2L, 3L, 
1L, 2L, 3L, 1L, 2L, 3L, 1L, 2L, 3L, 1L, 2L, 3L, 1L, 2L, 3L, 1L, 
2L, 3L, 1L, 2L, 3L, 1L, 2L, 3L, 1L, 2L), levels = c("1", "2", 
"3"), class = "factor"), C = c("0", "0", "0", "0", "0", "0", 
"1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "1", 
"1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "1", 
"1", "1", "1", "1", "0", "0", "0", "0", "0", "0", "1", "1", "1", 
"1", "1", "1", "0", "0"), D = c(1016.66666666667, 1016.66666666667, 
1018.05555555556, 1016.66666666667, 1016.66666666667, 1015.27777777778, 
1017.36111111111, 1017.36111111111, 1016.66666666667, 1015.97222222222, 
1015.97222222222, 1016.66666666667, 1018.05555555556, 1016.66666666667, 
1016.66666666667, 1015.27777777778, 1016.66666666667, 1016.66666666667, 
1017.36111111111, 1017.36111111111, 1019.44444444444, 1015.97222222222, 
1015.97222222222, 1013.88888888889, 1017.36111111111, 1018.05555555556, 
1017.36111111111, 1015.97222222222, 1015.27777777778, 1015.97222222222, 
1018.05555555556, 1017.36111111111, 1016.66666666667, 1015.27777777778, 
1015.97222222222, 1016.66666666667, 1016.66666666667, 1016.66666666667, 
1016.66666666667, 1016.66666666667, 1016.66666666667, 1016.66666666667, 
1016.66666666667, 1016.66666666667, 1016.66666666667, 1016.66666666667, 
1016.66666666667, 1016.66666666667, 1021.52777777778, 1018.05555555556
)), row.names = c(NA, 50L), class = "data.frame")

错误原因

报错核心是数据存在重复的被试-处理组合:

  • rstatix::anova_test执行时会将长格式数据转为宽格式,要求每个被试(wid=A)在每个处理组合(within=B + between=C)下仅有一个观测值。
  • 查看示例数据:比如被试A="1"、C="0"时,B的每个水平(1、2、3)都对应两行数据,这种重复组合导致spread()无法生成唯一行,从而抛出错误。

解决方法

方法1:删除重复观测值

如果重复是数据录入错误,直接保留每个(A,B,C)组合的唯一记录:

library(dplyr)
# 去重,保留每个组合的第一行数据
data_ex_clean <- data_ex %>% distinct(A, B, C, .keep_all = TRUE)

方法2:合并重复观测值(取均值)

如果重复是同一条件下的多次测量,先计算每个条件下的均值:

library(dplyr)
# 按被试和处理组合分组,计算均值
data_ex_summary <- data_ex %>% 
  group_by(A, B, C) %>% 
  summarise(D = mean(D, na.rm = TRUE), .groups = "drop")

重新运行ANOVA

用清理后的数据执行原代码:

res.aov = anova_test(data_ex_clean, # 或替换为data_ex_summary
                     dv = D, 
                     wid = A, 
                     within = B,
                     between = C,
                     type = 3)

内容的提问来源于stack exchange,提问作者12666727b9

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 19:27:02