You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何调整R语言配对Wilcoxon检验结果的输出格式?

调整R语言配对Wilcoxon检验结果输出格式

我编写了一段R语言代码,基于指定的列组合执行配对Wilcoxon检验,计算每组的中位数(med)与p值,但当前输出格式不符合需求。现提供原始数据、现有函数代码及期望的输出数据框格式,请求调整代码以得到目标格式的结果。

原始数据

data <- structure(list(col1 = 1:9, col2 = 10:18, col3 = 16:24, col4 = 67:75, 
    col5 = c(19L, 19L, 19L, 19L, 19L, 19L, 19L, 19L, 19L), GROUP = c(1L, 
    1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L)), class = "data.frame", row.names = c(NA, 
-9L))

现有代码

combination <- list(c(1, 2), c(1, 3), c(1, 4),c(2,3),c(2,4),c(2,5),c(3,4),c(3,5))

wilcox.fun <- function(dat) { 
  do.call(rbind, lapply(combination, function(x) {
    test <- wilcox.test(dat[[x[1]]], dat[[x[2]]], paired=TRUE)
    data.frame(Test = sprintf('%s by %s', x[1],x[2]), 
               #W = round(test$statistic,4),
               med = paste(median(dat[[x[1]]]),median(dat[[x[2]]])),
               p = test$p.value)
  }))
}

result <- purrr::map_df(split(data, data$GROUP), wilcox.fun, .id = 'Group')

期望输出格式

resulte <- structure(list(Group = c(1L, 1L, 1L, 1L, 1L), Test = 1:5, med = c(5L, 
14L, 20L, 71L, 19L), p = c("1 by 2: 0,00335343645494632\n1 by 3: 0,00335343645494632\n1 by 4: 0,00335343645494632", 
"2 by 3: 0,00335343645494632\n2 by 4: 0,00335343645494632\n2 by 5: 0,00390625\n", 
"3 by 4: 0,00335343645494632\n3 by 5: 0,325204163250902", NA, 
NA)), class = "data.frame", row.names = c(NA, -5L))

调整后的代码

library(purrr)

combination <- list(c(1, 2), c(1, 3), c(1, 4),c(2,3),c(2,4),c(2,5),c(3,4),c(3,5))

wilcox.fun <- function(dat) {
  # 提取前5列的中位数,对应Test1-5的med值
  col_medians <- sapply(dat[, 1:5], median)
  
  # 执行所有配对检验,整理检验标签和p值
  test_details <- lapply(combination, function(x) {
    test_res <- wilcox.test(dat[[x[1]]], dat[[x[2]]], paired = TRUE)
    list(main_col = x[1],
         test_text = sprintf('%s by %s: %s', x[1], x[2], test_res$p.value))
  })
  
  # 按主列分组,合并同主列的检验结果为换行字符串
  grouped_p <- split(test_details, sapply(test_details, function(x) x$main_col))
  p_combined <- lapply(grouped_p, function(g) paste(sapply(g, function(x) x$test_text), collapse = '\n'))
  
  # 构建目标数据框
  data.frame(
    Group = unique(dat$GROUP),
    Test = 1:5,
    med = col_medians,
    p = c(p_combined[as.character(1:3)], NA, NA),
    stringsAsFactors = FALSE
  )
}

# 生成最终结果
result <- purrr::map_df(split(data, data$GROUP), wilcox.fun)

代码说明

  1. 先提取数据前5列的中位数,直接对应期望输出中med列的数值;
  2. 批量执行所有指定列组合的配对Wilcoxon检验,记录每个检验的主列、完整标签和p值;
  3. 按主列(x by y中的x)分组,将同主列的检验结果合并为换行分隔的字符串;
  4. 构建符合要求的数据框,将合并后的p值对应到Test1-3,Test4和5的p值设为NA,完全匹配期望输出格式。

内容的提问来源于stack exchange,提问作者GOGA GOGA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 15:45:50