You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何生成以手术指征为行、患者占比为统计量的gtsummary表格

解决gtsummary统计手术指征时的结果不符合预期问题

问题原因

你之前将长格式数据转成宽格式后,dcast生成的列值是每位患者对应指征的出现次数(比如患者3的infection列值为2,代表该患者有2次感染指征)。而gtsummary::tbl_summary默认会对这类数值型变量统计其取值分布(比如统计infection列中0、1、2各自出现的患者数),这和你需要的「统计每个指征的总事件数、有该指征的患者数及占比」完全不符,所以得到了错误的输出。

解决方案

不需要转宽格式,直接基于原始长格式数据处理即可。以下提供两种可行方案:

方案一:先手动汇总统计,再生成表格

先按手术指征分组,计算每个指征的总事件数、对应患者数及占比,再用gtsummary生成表格:

library(data.table)
library(gtsummary)
library(stringr)

# 原始数据
participant.index <- c(1, 1, 2, 3, 3, 4, 5, 5, 6, 7, 7)
indication.surgery <- c(
  "ibs",
  "infection", 
  "infection",
  "infection",
  "infection",
  "ibs",
  "infection",
  "infection",
  "ibs",
  "tumour",
  "tumour"
)

mydata <- data.table(participant.index, indication.surgery)

# 计算总患者数
total_patients <- uniqueN(mydata$participant.index)

# 按指征分组统计核心指标
summary_data <- mydata[, .(
  n = .N,  # 总事件数
  patient_n = uniqueN(participant.index),  # 有该指征的患者数
  patient_pct = uniqueN(participant.index)/total_patients  # 患者占比
), by = indication.surgery]

# 整理格式,符合期望输出的列名和样式
summary_data <- summary_data[, .(
  `Indication for surgery` = str_to_title(indication.surgery),
  n = n,
  `N (%)` = sprintf("%d/%d (%.2f%%)", patient_n, total_patients, patient_pct*100)
)]

# 生成gtsummary表格
final_table <- summary_data %>%
  tbl_summary(
    include = c(`Indication for surgery`, n, `N (%)`),
    type = all_continuous() ~ "continuous2",
    statistic = all_continuous() ~ "{mean}",  # 因为n和N(%)是固定值,用mean直接展示
    label = list(n ~ "n", `N (%)` ~ "N (%)")
  ) %>%
  modify_header(
    label ~ "**Indication for surgery**",
    stat_0 ~ "**n**",
    stat_1 ~ "**N (%)**"
  ) %>%
  modify_column_alignment(columns = everything(), align = "center")

final_table

方案二:自定义gtsummary统计函数,直接处理原始数据

通过自定义统计函数,让tbl_summary直接输出你需要的指标:

library(data.table)
library(gtsummary)
library(dplyr)
library(stringr)

# 原始数据
participant.index <- c(1, 1, 2, 3, 3, 4, 5, 5, 6, 7, 7)
indication.surgery <- c(
  "ibs",
  "infection", 
  "infection",
  "infection",
  "infection",
  "ibs",
  "infection",
  "infection",
  "ibs",
  "tumour",
  "tumour"
)

mydata <- data.table(participant.index, indication.surgery) %>%
  mutate(indication.surgery = str_to_title(indication.surgery))

total_patients <- n_distinct(mydata$participant.index)

# 自定义统计函数:返回每个指征的事件数、患者数及占比
custom_stat <- function(data, variable, ...) {
  group_data <- data %>%
    group_by({{variable}}) %>%
    summarize(
      event_count = n(),
      patient_count = n_distinct(participant.index),
      patient_ratio = patient_count/total_patients
    )
  
  # 整理成gtsummary要求的输出格式
  tibble(
    label = group_data[[variable]],
    stat_0 = group_data$event_count,
    stat_1 = sprintf("%d/%d (%.2f%%)", group_data$patient_count, total_patients, group_data$patient_ratio*100)
  )
}

# 生成表格
final_table <- mydata %>%
  tbl_summary(
    include = indication.surgery,
    type = indication.surgery ~ "categorical",
    statistic = indication.surgery ~ custom_stat
  ) %>%
  modify_header(
    label ~ "**Indication for surgery**",
    stat_0 ~ "**n**",
    stat_1 ~ "**N (%)**"
  ) %>%
  modify_column_alignment(columns = everything(), align = "center")

final_table

两种方案都能生成符合你预期的表格,其中方案一更直观,方案二则更贴合gtsummary的使用逻辑。

内容的提问来源于stack exchange,提问作者Shashi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 06:30:05