You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在tbl_summary中排除部分因子水平计算占比与统计检验?

实现方法

要达成需求,我们可以通过自定义统计量函数和自定义检验函数,结合gtsummary包的功能来实现,具体步骤如下:

1. 加载所需包

首先确保安装并加载相关R包:

library(data.table)
library(gtsummary)
library(dplyr)
library(glue)

2. 定义排除类别与自定义函数

先指定需要排除的收入类别,然后分别定义用于计算百分比的统计量函数,以及用于生成p值的检验函数:

# 定义需要排除的类别
exclude_cats <- c("Not applicable", "I rather not to answer")

# 自定义统计量:排除指定类别后计算百分比
custom_stat <- function(x, variable, by, ...) {
  # 计算当前分组中有效样本量(排除指定类别)
  n_valid <- sum(!x %in% exclude_cats)
  # 为每个类别生成统计文本
  purrr::map_chr(unique(x), ~{
    n_count <- sum(x == .x)
    if (.x %in% exclude_cats) {
      # 排除类别仅显示计数
      glue("{n_count}")
    } else {
      # 有效类别显示计数+百分比(分母为有效样本量)
      pct <- round(n_count / n_valid * 100, 1)
      glue("{n_count} ({pct}%)")
    }
  })
}

# 自定义Fisher检验:仅用有效样本计算p值
custom_fisher <- function(data, variable, by, ...) {
  # 过滤掉排除类别的数据
  data_filtered <- data %>% filter(!.data[[variable]] %in% exclude_cats)
  # 执行Fisher精确检验
  test_result <- fisher.test(data_filtered[[by]], data_filtered[[variable]])
  # 返回符合gtsummary要求的结果格式
  tibble::tibble(p.value = test_result$p.value, method = "Fisher's exact test")
}

3. 生成统计表格

调用gtsummary的函数,应用自定义函数生成目标表格:

example %>% 
  tbl_summary(
    by = sex,
    include = income,
    # 为income列指定自定义统计量
    statistic = list(income ~ custom_stat),
    # 保留排除类别,不将其标记为缺失值隐藏
    missing = "no"
  ) %>%
  # 添加自定义检验的p值
  add_p(test = list(income ~ custom_fisher)) %>%
  # 可选:添加总样本列,优化表格样式
  add_overall() %>%
  modify_header(label = "**Income Category**") %>%
  bold_labels()

效果说明

  • 表格中"Not applicable"和"I rather not to answer"仅显示计数,不参与其他类别的百分比计算(百分比分母为对应性别中排除这两类后的有效样本量)。
  • p值完全来自过滤掉排除类别后的fisher.test结果。

内容的提问来源于stack exchange,提问作者MDSF

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 01:22:30