如何用gtsummary在R中按两类变量汇总分类数据并生成P值
解决方案
可以用R的dplyr、tidyr和broom包实现需求,步骤如下:
1. 加载依赖包并准备数据
library(dplyr) library(tidyr) library(broom) library(scales) # 示例数据 df <- structure(list(success = c(1, 1, 1, 1, 1), stent_type = c(1, 1, 1, 0, 1), stent_30 = c(TRUE, TRUE, FALSE, FALSE, TRUE)), row.names = c(NA, 5L), class = "data.frame")
2. 分组计算成功数与百分比
按stent_type和stent_30分组,统计每组成功案例数、总案例数,并计算成功率:
summary_data <- df %>% mutate(stent_30_label = case_when( stent_30 == TRUE ~ "支架放置<30天", TRUE ~ "支架放置>30天" )) %>% group_by(stent_type, stent_30_label) %>% summarise( success_count = sum(success), total_count = n(), success_pct = percent(success_count / total_count, accuracy = 1), .groups = "drop" ) %>% mutate(display_text = paste0(success_pct, " (", success_count, ")"))
3. 计算每个stent_type分组的P值
针对每个stent_type,对比stent_30两组的成功率差异,用Fisher精确检验(样本量小时更合适):
p_value_data <- df %>% group_by(stent_type) %>% summarise( p_value = fisher.test(table(stent_30, success))$p.value, .groups = "drop" ) %>% mutate(p_value = round(p_value, 2)) # 保留两位小数
4. 合并数据并整理成目标表格
把汇总数据转成宽格式,再和P值数据合并:
final_table <- summary_data %>% pivot_wider( id_cols = stent_type, names_from = stent_30_label, values_from = display_text ) %>% left_join(p_value_data, by = "stent_type") %>% rename( " " = stent_type, "P值" = p_value ) %>% mutate(` ` = paste0("stent ", ` `))
运行后final_table即为目标格式,示例数据输出结果:
| 支架放置<30天 | 支架放置>30天 | P值 | |
|---|---|---|---|
| stent 0 | 100% (1) | 100% (1) | 1.00 |
| stent 1 | 100% (3) | 100% (1) | 1.00 |
补充说明
- 若样本量较大,可替换
fisher.test()为chisq.test()计算P值 scales::percent()用于格式化百分比,accuracy=1表示保留整数百分比
内容的提问来源于stack exchange,提问作者jon
相关产品推荐
相关产品推荐

