You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

宽格式下summarise调用wilcox.test报错的原因与解决方法

问题描述

我有一个包含Status字段(仅包含PRE、POST两个水平)的数据框,希望按列比较不同Status水平下的数值差异,使用wilcoxon秩和检验。尝试用宽格式代码执行时报错,但转为长格式后代码可正常运行。想了解宽格式下代码执行失败的原因,以及如何在宽格式下实现需求。

附数据结构:

df=structure(list(SI_mean = c(2.6, 3, 2.4, 3, 3, 3.2, 2.2, 4, 3.8, 
2.8, 3.6, 2, 3.6, 3.6, 3.8, 3.2, 3, 4, 4, 3, 3.2, 4, 4, 3.2, 
3.2, 3, 3.2, 3.8, 4, 4, 4, 3), TU_mean = c(2.75, 3, 2.75, 3, 
3, 2.75, 3, 3.5, 3.75, 2.5, 3.25, 2, 3.5, 4, 3, 3.25, 3, 4, 4, 
3, 4, 4, 4, 3.25, 3.25, 3, 3, 3.25, 4, 4, 4, 3), ED_mean = c(2.6, 
2.4, 2.4, 3, 2.8, 4, 2, 3.8, 2.6, 2, 2.8, 2, 3, 3.4, 3, 1, 3, 
4, 3.8, 3, 3, 4, 4, 3.2, 4, 2.6, 4, 4, 3.8, 3.6, 4, 3), MT_mean = c(2.8, 
2.4, 2.2, 3, 2.8, 3.4, 2.2, 3.6, 3.4, 3, 2.6, 1.8, 3.4, 3, 4, 
2, 3, 3.4, 3.4, 3, 4, 4, 4, 3.2, 4, 2.8, 4, 4, 3.8, 3.6, 4, 3
), DT_mean = c(3.4, 3, 2.6, 3, 3, 3.8, 2.4, 3.6, 3, 3, 2.8, 2.4, 
3.6, 3.6, 3, 2.2, 3, 4, 4, 4, 3.6, 4, 4, 3.6, 3.8, 2.8, 4, 4, 
4, 3.8, 4, 3), SK_mean = c(2.5, 3, 2.25, 3, 3, 3.5, 2.25, 4, 
3.25, 2.25, 2.5, 2.5, 3.75, 3.75, 4, 1, 2, 4, 3.25, 3, 3.75, 
4, 4, 2.75, 4, 3, 4, 4, 4, 4, 4, 3), ATT_mean = c(3.8, 4, 2.8, 
3, 3, 3.8, 3, 3.6, 3, 4, 4, 3, 3.8, 3.6, 4, 3.8, 4, 4, 4, 4, 
4, 4, 4, 3.6, 3.8, 3, 4, 4, 4, 4, 4, 4), Status = c("PRE", "PRE", 
"PRE", "PRE", "PRE", "PRE", "PRE", "PRE", "PRE", "PRE", "PRE", 
"PRE", "PRE", "PRE", "PRE", "PRE", "PRE", "POST", "POST", "POST", 
"POST", "POST", "POST", "POST", "POST", "POST", "POST", "POST", 
"POST", "POST", "POST", "POST")), class = c("rowwise_df", "tbl_df", 
"tbl", "data.frame"), row.names = c(NA, -32L), groups = structure(list(
    .rows = structure(list(1L, 2L, 3L, 4L, 5L, 6L, 7L, 8L, 9L, 
        10L, 11L, 12L, 13L, 14L, 15L, 16L, 17L, 18L, 19L, 20L, 
        21L, 22L, 23L, 24L, 25L, 26L, 27L, 28L, 29L, 30L, 31L, 
        32L), ptype = integer(0), class = c("vctrs_list_of", 
    "vctrs_vctr", "list"))), row.names = c(NA, -32L), class = c("tbl_df", 
"tbl", "data.frame")))

报错的宽格式代码

df |>
  summarise(across(contains("mean"),~wilcox.test(.x~Status)$p.value))

错误信息

Error in `summarise()`:
ℹ In argument: `across(contains("mean"), ~wilcox.test(.x ~ Status)$p.value)`.
ℹ In row 1.
Caused by error in `across()`:
! Can't compute column `SI_mean`.
Caused by error in `wilcox.test.formula()`:
! grouping factor must have exactly 2 levels

正常运行的长格式代码

df |> pivot_longer(contains("mean"),names_to = "Variable",values_to = "Mean") |>
  group_by(Variable) |>
  summarise(
    wilcox_p_value = wilcox.test(Mean~Status)$p.value
    )

宽格式代码失败的原因

你的数据框df是**行分组(rowwise_df)**格式(从数据结构的class = c("rowwise_df", ...)可看出):

  • 在rowwise分组模式下,summarise会逐行处理数据,每一行中的Status只有单个值(要么是PRE,要么是POST),此时调用wilcox.test(.x~Status)时,分组因子Status仅存在1个水平,不符合wilcoxon检验需要2个分组水平的要求,因此触发报错。
  • 长格式代码中,pivot_longer会自动取消rowwise分组,转为常规数据框,之后按Variable分组时,每个组内包含了PRE和POST两个水平的Status,检验可以正常执行。

宽格式下的实现方法

只需先取消rowwise分组,再执行检验即可:

方法1:取消行分组后执行检验

这是最直接且推荐的方式:

df |>
  ungroup() |>  # 关键步骤:取消行分组
  summarise(across(contains("mean"), 
                   ~wilcox.test(.x ~ Status, data = cur_data())$p.value))

方法2:手动传入完整数据(不推荐)

如果不想取消分组,可以在wilcox.test中手动传入整个数据框作为data参数,确保能获取所有行的Status和对应变量值:

df |>
  summarise(across(contains("mean"), 
                   ~wilcox.test(reformulate("Status", response = cur_column()), 
                                data = df)$p.value))

结果验证

两种宽格式方法得到的结果与长格式代码一致,示例输出:

# A tibble: 1 × 7
  SI_mean TU_mean ED_mean MT_mean DT_mean SK_mean ATT_mean
    <dbl>   <dbl>   <dbl>   <dbl>   <dbl>   <dbl>    <dbl>
1  0.0819  0.0819  0.0455  0.0455  0.136   0.0455   0.0819

内容的提问来源于stack exchange,提问作者Derf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 17:04:56