You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在gtsummary::tbl_summary()中展示by变量NA计数且排除在行百分比分母外

问题

在gtsummary::tbl_summary()中,希望实现以下需求:

  • 展示分组(by参数指定)变量的缺失值(NA)计数,但不将这些NA纳入行百分比的分母
  • 仅在“缺失值”列显示计数,且单独为该列切换为列百分比(其余非缺失列使用行百分比)
    目前通过tbl_merge()实现了备选方案,想知道是否有更直接的解决方法?

原示例代码:

library(gtsummary)

tbl1 <- 
  trial |>
    tbl_summary(
        by = response, 
        include = c(age, grade, trt), 
        type = list(
          all_dichotomous() ~ "categorical",
          age ~ "continuous"),
        missing = "no",
        percent = "row"
    ) |>
    add_p()

tbl2 <- 
  trial |>
  filter(is.na(response)) |>
    tbl_summary(
        # by = response, 
        include = c(age, grade, trt), 
        type = list(
          all_dichotomous() ~ "categorical",
          age ~ "continuous"),
        missing = "no",
        percent = "column"
    )

list(tbl1, tbl2) |>
  tbl_merge(tab_spanner = c("**Tumor Response**", "**Missing**"))

原输出效果:
原代码输出效果

解决方法

这里提供两种更简洁的实现方式:

方法一:将缺失值转为显式分组(无需拆分表格)

直接把response的缺失值转为一个单独分组,再调整不同分组的百分比计算规则:

library(gtsummary)
library(dplyr)
library(forcats)
library(stringr)

# 把response的缺失值转为显式分组
trial_mod <- trial |> 
  mutate(response = fct_explicit_na(response, na_level = "Missing"))

# 生成汇总表并调整百分比规则
tbl_final <- trial_mod |>
  tbl_summary(
    by = response,
    include = c(age, grade, trt),
    type = list(all_dichotomous() ~ "categorical", age ~ "continuous"),
    missing = "no",
    percent = "column"  # 先统一用列百分比,后续调整非缺失组为行百分比
  ) |>
  add_p() |>
  # 调整非Missing组的百分比为行百分比(分母仅统计非缺失样本)
  modify_table_body(
    ~ .x |>
      mutate(
        across(
          starts_with("stat_") & !ends_with("Missing"),
          ~ case_when(
            variable_type == "categorical" ~ {
              # 提取计数
              n_val <- as.numeric(str_extract(.x, "\\d+(?=\\s)"))
              # 计算非缺失样本中该变量的总例数
              row_total <- trial |> filter(!is.na(response)) |> pull(!!sym(variable)) |> table() |> sum()
              # 计算行百分比
              pct_val <- n_val / row_total * 100
              paste0(n_val, " (", sprintf("%.1f", pct_val), "%)")
            },
            # 连续变量保持原统计结果
            variable_type == "continuous" ~ .x
          )
        )
      )
  ) |>
  # 设置表头分栏
  modify_spanning_header(
    all_stat_cols() ~ case_when(
      colname == "stat_Missing" ~ "**Missing**",
      TRUE ~ "**Tumor Response**"
    )
  )

方法二:优化原tbl_merge思路(减少代码重复)

保留拆分合并的逻辑,但复用参数减少冗余:

library(gtsummary)

# 定义统一的汇总参数,避免重复编写
summary_params <- list(
  include = c(age, grade, trt),
  type = list(all_dichotomous() ~ "categorical", age ~ "continuous"),
  missing = "no"
)

# 生成非缺失分组的主表(行百分比)
tbl_main <- trial |>
  filter(!is.na(response)) |>
  tbl_summary(
    by = response,
    percent = "row",
    !!!summary_params  # 复用参数
  ) |>
  add_p()

# 生成缺失值列的子表(列百分比)
tbl_missing <- trial |>
  filter(is.na(response)) |>
  tbl_summary(
    percent = "column",
    !!!summary_params  # 复用参数
  )

# 合并表格
tbl_final <- list(tbl_main, tbl_missing) |>
  tbl_merge(tab_spanner = c("**Tumor Response**", "**Missing**"))

说明

  • 方法一无需拆分表格,通过显式分组+手动调整百分比实现需求,逻辑连贯
  • 方法二优化了原代码的冗余问题,用!!!运算符复用参数,代码更简洁易维护

内容的提问来源于stack exchange,提问作者Doug Ezra Morrison

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 17:29:54