You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用anova_test执行重复测量ANOVA时spread()报错:如何让键唯一?

混合设计ANOVA报错解决:spread()键重复问题

问题重现

尝试分析condition(组间)与time(组内)对收缩压(syst)的影响时,使用rstatix::anova_test()出现如下错误:

Error in `spread()`:
! Each row of output must be identified by a unique combination of keys.
ℹ Keys are shared for 6 rows
• 14, 36
• 182, 204
• 98, 120

涉及的样本数据如下:

> df_lbp[c(14,36),]
# A tibble: 2 × 10
  subject_id condition visit  time  syst  timef      conditionf
       <dbl>     <dbl> <dbl> <int> <dbl>  <fct>      <fct>     
1        129         0     1     1  106.  anticipate Control   
2        165         1     1     1  119   anticipate Stress    
> df_lbp[c(182, 204),]
# A tibble: 2 × 10
  subject_id condition visit  time  syst  timef    conditionf
       <dbl>     <dbl> <dbl> <int> <dbl>  <fct>    <fct>     
1        129         0     1     3  103.  recovery Control   
2        165         1     1     3  121.  recovery Stress    
> df_lbp[c(98, 120),]
# A tibble: 2 × 10
  subject_id condition visit  time  syst  timef conditionf
       <dbl>     <dbl> <dbl> <int> <dbl>  <fct> <fct>     
1        129         0     1     2  102.  task  Control   
2        165         1     1     2  128   task  Stress 

原代码:

a1 <- anova_test( data = df_lbp, dv = syst,
                  wid = subject_id, 
                  within = c(timef, conditionf) )

get_anova_table(a1)

问题原因

  1. 变量类型误用:
    conditionf是组间变量(每个受试者仅属于Control或Stress中的一组),但被错误地放入within参数中。anova_test()的within参数要求变量是组内重复测量变量(即每个受试者在该变量的所有水平下都有观测值),这种参数设置会导致函数逻辑冲突。
  2. 键的定义:
    函数内部会将wid(subject_id)与所有within变量的组合作为唯一键,要求每个键对应唯一一行数据。由于conditionf是组间变量,每个受试者仅对应其水平,函数在转换数据格式时会触发键重复的错误提示。

解决步骤

1. 修正ANOVA参数(混合设计)

将组间变量放入between参数,组内变量保留在within参数,对应混合设计ANOVA:

a1 <- anova_test(
  data = df_lbp, 
  dv = syst,
  wid = subject_id, 
  between = conditionf,  # 组间变量:Control/Stress
  within = timef         # 组内变量:三个时间点
)
get_anova_table(a1)

2. 检查并清理重复数据

确保每个受试者在对应变量组合下没有重复观测,执行以下代码排查:

library(dplyr)
# 检查subject_id + conditionf + timef的重复项
df_lbp %>%
  count(subject_id, conditionf, timef) %>%
  filter(n > 1)

若输出结果不为空,说明存在重复行,可使用distinct()删除重复:

df_lbp_clean <- df_lbp %>% distinct(subject_id, conditionf, timef, .keep_all = TRUE)

再用清理后的数据重新运行ANOVA。

内容的提问来源于stack exchange,提问作者Mia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 01:42:42