如何调整R语言配对Wilcoxon检验结果的输出格式?
调整R语言配对Wilcoxon检验结果输出格式
我编写了一段R语言代码,基于指定的列组合执行配对Wilcoxon检验,计算每组的中位数(med)与p值,但当前输出格式不符合需求。现提供原始数据、现有函数代码及期望的输出数据框格式,请求调整代码以得到目标格式的结果。
原始数据
data <- structure(list(col1 = 1:9, col2 = 10:18, col3 = 16:24, col4 = 67:75, col5 = c(19L, 19L, 19L, 19L, 19L, 19L, 19L, 19L, 19L), GROUP = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L)), class = "data.frame", row.names = c(NA, -9L))
现有代码
combination <- list(c(1, 2), c(1, 3), c(1, 4),c(2,3),c(2,4),c(2,5),c(3,4),c(3,5)) wilcox.fun <- function(dat) { do.call(rbind, lapply(combination, function(x) { test <- wilcox.test(dat[[x[1]]], dat[[x[2]]], paired=TRUE) data.frame(Test = sprintf('%s by %s', x[1],x[2]), #W = round(test$statistic,4), med = paste(median(dat[[x[1]]]),median(dat[[x[2]]])), p = test$p.value) })) } result <- purrr::map_df(split(data, data$GROUP), wilcox.fun, .id = 'Group')
期望输出格式
resulte <- structure(list(Group = c(1L, 1L, 1L, 1L, 1L), Test = 1:5, med = c(5L, 14L, 20L, 71L, 19L), p = c("1 by 2: 0,00335343645494632\n1 by 3: 0,00335343645494632\n1 by 4: 0,00335343645494632", "2 by 3: 0,00335343645494632\n2 by 4: 0,00335343645494632\n2 by 5: 0,00390625\n", "3 by 4: 0,00335343645494632\n3 by 5: 0,325204163250902", NA, NA)), class = "data.frame", row.names = c(NA, -5L))
调整后的代码
library(purrr) combination <- list(c(1, 2), c(1, 3), c(1, 4),c(2,3),c(2,4),c(2,5),c(3,4),c(3,5)) wilcox.fun <- function(dat) { # 提取前5列的中位数,对应Test1-5的med值 col_medians <- sapply(dat[, 1:5], median) # 执行所有配对检验,整理检验标签和p值 test_details <- lapply(combination, function(x) { test_res <- wilcox.test(dat[[x[1]]], dat[[x[2]]], paired = TRUE) list(main_col = x[1], test_text = sprintf('%s by %s: %s', x[1], x[2], test_res$p.value)) }) # 按主列分组,合并同主列的检验结果为换行字符串 grouped_p <- split(test_details, sapply(test_details, function(x) x$main_col)) p_combined <- lapply(grouped_p, function(g) paste(sapply(g, function(x) x$test_text), collapse = '\n')) # 构建目标数据框 data.frame( Group = unique(dat$GROUP), Test = 1:5, med = col_medians, p = c(p_combined[as.character(1:3)], NA, NA), stringsAsFactors = FALSE ) } # 生成最终结果 result <- purrr::map_df(split(data, data$GROUP), wilcox.fun)
代码说明
- 先提取数据前5列的中位数,直接对应期望输出中
med列的数值; - 批量执行所有指定列组合的配对Wilcoxon检验,记录每个检验的主列、完整标签和p值;
- 按主列(
x by y中的x)分组,将同主列的检验结果合并为换行分隔的字符串; - 构建符合要求的数据框,将合并后的p值对应到Test1-3,Test4和5的p值设为NA,完全匹配期望输出格式。
内容的提问来源于stack exchange,提问作者GOGA GOGA
相关产品推荐
相关产品推荐

