如何在tbl_summary中排除部分因子水平计算占比与统计检验?
实现方法
要达成需求,我们可以通过自定义统计量函数和自定义检验函数,结合gtsummary包的功能来实现,具体步骤如下:
1. 加载所需包
首先确保安装并加载相关R包:
library(data.table) library(gtsummary) library(dplyr) library(glue)
2. 定义排除类别与自定义函数
先指定需要排除的收入类别,然后分别定义用于计算百分比的统计量函数,以及用于生成p值的检验函数:
# 定义需要排除的类别 exclude_cats <- c("Not applicable", "I rather not to answer") # 自定义统计量:排除指定类别后计算百分比 custom_stat <- function(x, variable, by, ...) { # 计算当前分组中有效样本量(排除指定类别) n_valid <- sum(!x %in% exclude_cats) # 为每个类别生成统计文本 purrr::map_chr(unique(x), ~{ n_count <- sum(x == .x) if (.x %in% exclude_cats) { # 排除类别仅显示计数 glue("{n_count}") } else { # 有效类别显示计数+百分比(分母为有效样本量) pct <- round(n_count / n_valid * 100, 1) glue("{n_count} ({pct}%)") } }) } # 自定义Fisher检验:仅用有效样本计算p值 custom_fisher <- function(data, variable, by, ...) { # 过滤掉排除类别的数据 data_filtered <- data %>% filter(!.data[[variable]] %in% exclude_cats) # 执行Fisher精确检验 test_result <- fisher.test(data_filtered[[by]], data_filtered[[variable]]) # 返回符合gtsummary要求的结果格式 tibble::tibble(p.value = test_result$p.value, method = "Fisher's exact test") }
3. 生成统计表格
调用gtsummary的函数,应用自定义函数生成目标表格:
example %>% tbl_summary( by = sex, include = income, # 为income列指定自定义统计量 statistic = list(income ~ custom_stat), # 保留排除类别,不将其标记为缺失值隐藏 missing = "no" ) %>% # 添加自定义检验的p值 add_p(test = list(income ~ custom_fisher)) %>% # 可选:添加总样本列,优化表格样式 add_overall() %>% modify_header(label = "**Income Category**") %>% bold_labels()
效果说明
- 表格中"Not applicable"和"I rather not to answer"仅显示计数,不参与其他类别的百分比计算(百分比分母为对应性别中排除这两类后的有效样本量)。
- p值完全来自过滤掉排除类别后的
fisher.test结果。
内容的提问来源于stack exchange,提问作者MDSF
相关产品推荐
相关产品推荐

