如何为gtsummary表格按病例数分配行权重?含frequency_weights咨询
gtsummary加权对比表格解决方案
问题1:按病例数分配权重的简便方法
无需转换为长格式,直接在tbl_summary()中结合权重参数即可实现按no_patients加权计算。因为你的数据是研究级聚合数据,每行对应no_patients个患者样本,用频率权重就能让统计结果自动向样本量大的研究倾斜,完全适配多变量复杂数据集。
示例代码:
library(tidyverse) library(gtsummary) data <- data.frame( study_id = c(1, 2, 3, 4, 5), no_patients = c(10, 15, 20, 23, 16), treatment_id = c("surgery", "radiotherapy", "surgery", "radiotherapy", "surgery"), treatment_outcome = c(0.88, 0.50, 0.90, 0.23, 0.67), recurrence = c(0, 2, 4, 3, 6) ) # 生成带病例数权重的对比表格 weighted_tbl <- data %>% tbl_summary( by = treatment_id, weights = frequency_weights(no_patients), # 指定病例数为频率权重 include = c(treatment_outcome, recurrence), statistic = list( treatment_outcome ~ "{mean}", # 计算加权平均治疗结局 recurrence ~ "{mean}" # 计算加权平均复发次数 ) ) %>% add_p() # 可选,组间比较同样会应用权重 weighted_tbl
问题2:关于frequency_weights()工具
frequency_weights()是gtsummary包专为频率权重设计的工具,它的作用是告知统计函数:每行数据代表了weights参数指定数量的个体。你的场景中,每行对应一个研究的no_patients个患者,正好匹配频率权重的使用场景。
它和probability_weights()区分使用:后者适用于权重代表抽样概率的场景,而你的研究级聚合数据用频率权重完全合适。只需将其传入tbl_summary()的weights参数,后续所有统计计算(描述统计、组间比较)都会自动应用该权重。
内容的提问来源于stack exchange,提问作者JLA
相关产品推荐
相关产品推荐

