如何在gtsummary包中正确计算配对Cohen's d?
解决gtsummary中计算配对Cohen's d的问题
错误原因
你遇到的报错cannot use 'paired = TRUE' in formula method是因为effectsize::cohens_d的公式调用方式不支持paired=TRUE参数,必须改用向量输入或宽格式数据来计算配对效应量。
正确步骤与代码
首先需要为配对数据添加个体ID(否则无法识别配对关系),然后自定义配对Cohen's d的计算函数,再传入add_difference中:
- 准备带配对ID的数据
time <- rep(c("1", "2"), each = 10) A <- 20:39 # 无需额外rep,20:39已生成20个连续值 B <- 100:119 id <- rep(1:10, 2) # 每个ID对应1次time1和1次time2的观测 df <- data.frame(id = id, time = time, A = A, B = B)
- 自定义配对Cohen's d计算函数
paired_cohens_d <- function(data, variable, by, ...) { # 将长格式数据转为宽格式,按ID分组 wide_data <- tidyr::pivot_wider( data, id_cols = id, names_from = all_of(by), values_from = all_of(variable) ) # 用向量方式调用cohens_d,指定paired=TRUE es_result <- effectsize::cohens_d( x = wide_data[[2]], y = wide_data[[3]], paired = TRUE, ci = 0.95 ) # 返回gtsummary要求的格式:estimate、置信区间、方法说明 tibble::tibble( estimate = es_result$Cohens_d, conf.low = es_result$CI_low, conf.high = es_result$CI_high, method = "配对Cohen's d" ) }
- 结合gtsummary生成表格并计算配对效应量
library(gtsummary) library(tidyr) df %>% tbl_summary( by = time, statistic = list(all_continuous() ~ "{mean} ({sd})") ) %>% add_difference( adj.vars = c(B), list(all_continuous() ~ paired_cohens_d) )
简化写法(直接在effectsize参数中定义匿名函数)
如果不想单独定义函数,也可以在add_difference的effectsize参数中直接写匿名函数:
df %>% tbl_summary( by = time, statistic = list(all_continuous() ~ "{mean} ({sd})") ) %>% add_difference( paired = TRUE, # 指定配对t检验(若需要同时输出t检验结果) adj.vars = c(B), effectsize = list(all_continuous() ~ function(data, variable, by, ...) { wide <- tidyr::pivot_wider(data, id_cols = id, names_from = by, values_from = variable) es <- effectsize::cohens_d(wide[[2]], wide[[3]], paired = TRUE, ci = 0.95) tibble::tibble(estimate = es$Cohens_d, conf.low = es$CI_low, conf.high = es$CI_high) }) )
内容的提问来源于stack exchange,提问作者atnplab
相关产品推荐
相关产品推荐

