You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让gtsummary以唯一ID数为分母统计多分类变量?

解决方法

要让gtsummary以唯一ID数作为百分比计算的分母,核心是跳过默认的按行统计,改为基于唯一ID的自定义统计。以下提供两种可行方案:

方案1:提前汇总统计数据(直观易理解)

先手动计算每个类别对应的唯一ID计数和百分比,再用gtsummary格式化表格:

library(gtsummary)
library(tidyverse)

# 原始数据
df <-
  tibble::tribble(
    ~id, ~category, ~success,
    1, "red;blue", TRUE,
    2, "red", TRUE,
    3, "blue", FALSE,
    4, "green", FALSE,
    5, "green;blue", TRUE
  )

# 计算各分组的唯一ID总数(作为百分比分母)
total_ids <- n_distinct(df$id)
false_ids <- n_distinct(df$id[df$success == FALSE])
true_ids <- n_distinct(df$id[df$success == TRUE])

# 拆分类别并统计每个类别的ID数与百分比
summary_data <- df %>%
  separate_rows(category, sep = ";") %>%
  group_by(category) %>%
  summarise(
    total = n_distinct(id),
    false = n_distinct(id[success == FALSE]),
    true = n_distinct(id[success == TRUE]),
    .groups = "drop"
  ) %>%
  mutate(
    total_pct = glue::glue("{total} ({round(total/total_ids*100, 0)}%)"),
    false_pct = glue::glue("{false} ({round(false/false_ids*100, 0)}%)"),
    true_pct = glue::glue("{true} ({round(true/true_ids*100, 0)}%)")
  )

# 生成目标表格
summary_data %>%
  select(category, total_pct, false_pct, true_pct) %>%
  rename(
    "类别" = category,
    "总计, N = 5" = total_pct,
    "FALSE, N = 2" = false_pct,
    "TRUE, N = 3" = true_pct
  ) %>%
  tbl_summary(
    include = everything(),
    type = all_categorical() ~ "categorical",
    statistic = all_categorical() ~ "{level}"
  ) %>%
  bold_labels() %>%
  as_kable()

方案2:在gtsummary流程中自定义统计量(更贴合gtsummary工作流)

直接在拆分后的数据集上,通过add_stat添加基于唯一ID的统计,并修正表头的N值:

library(gtsummary)
library(tidyverse)

df <-
  tibble::tribble(
    ~id, ~category, ~success,
    1, "red;blue", TRUE,
    2, "red", TRUE,
    3, "blue", FALSE,
    4, "green", FALSE,
    5, "green;blue", TRUE
  )

df_per_category <- df %>% separate_rows(category, sep = ";")

df_per_category %>%
  tbl_summary(
    include = category,
    by = success,
    label = list(category ~ "类别"),
    sort = all_categorical() ~ "frequency",
    statistic = all_categorical() ~ "{skip}" # 跳过默认统计
  ) %>%
  # 添加总计列的自定义统计(计数+百分比)
  add_stat(
    fns = ~n_distinct(.x$id[.x$category == .y$level]) %>%
      paste0(" (", round(./n_distinct(df$id)*100, 0), "%)"),
    location = "overall"
  ) %>%
  # 添加分组列的自定义统计
  add_stat(
    fns = ~n_distinct(.x$id[.x$category == .y$level & .x$success == .y$by]) %>%
      paste0(" (", round(./n_distinct(df$id[df$success == .y$by])*100, 0), "%)"),
    location = "by"
  ) %>%
  # 修改表头,显示正确的唯一ID数量
  modify_header(
    overall ~ "**总计**, N = 5",
    stat_1 ~ "**FALSE**, N = 2",
    stat_2 ~ "**TRUE**, N = 3"
  ) %>%
  bold_labels() %>%
  as_kable()

结果说明

两种方案都会生成你期望的表格:百分比以唯一ID数为分母,总和可超过100%,表头显示正确的分组ID总数。

内容的提问来源于stack exchange,提问作者maia-sh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 05:32:03