You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写有效函数批量格式化多个R语言gt表格?

批量格式化gt表格的有效实现方案

问题背景

需要对多个结构相似的数据框应用统一的gt表格格式,已实现单表格式化代码,但编写的批量格式化函数运行后未生效。

示例数据框

example <- data.frame(
  stringsAsFactors = FALSE,
               age = c("age 12 to 17", "age 18 to 25", "age 26 and up"),
           alc_est = c(14579.28558, 131872.35964, 660957.6512),
          bing_alc = c(0.0477, 0.3143, 0.224),
          marj_est = c(35913.3345, 137033.12968, 534667.52856),
            mj_use = c(0.1175, 0.3266, 0.1812),
        heroin_est = c(NA, 419.5748, 18294.36356),
        heroin_use = c(NA, 0.001, 0.0062),
          meth_est = c(550.16172, 2139.83148, 68161.25778),
          meth_use = c(0.0018, 0.0051, 0.0231)
)

原单表格式化代码

table_format <- example %>%
  gt(rowname_col = "age") %>%
  tab_stubhead(label = md("**Statewide**")) %>%
  tab_header(
    title = md("**table title here**"),
    subtitle = "more info here"
  ) %>%
  fmt_number(
      columns = c(2, 4, 6, 8),
      decimals = 0
  ) %>%
  tab_spanner_delim(
    delim = "_",
    columns = c(2:9)
  ) 

原未生效的批量函数

format_use_est <- function(df_here) {
  result <-  df_here %>%
    gt(rowname_col = "age") %>%
    tab_stubhead(label = md("**Statewide**")) %>%
    tab_header(
      title = md("**table title here**"),
      subtitle = "more info here"
    ) %>%
    fmt_percent(
      columns = c(3, 5, 7, 9),
      decimals = 1
    ) %>%
    fmt_number(
      columns = c(2, 4, 6, 8),
      decimals = 0
    ) %>%
    tab_spanner_delim(
      delim = "_",
      columns = c(2:9)
    ) 
} 

修复后的批量格式化函数

原函数未生效的核心原因是未显式返回格式化后的gt对象,同时硬编码列号的方式缺乏通用性。以下是优化后的函数:

format_use_est <- function(df_here, table_title = "**table title here**", table_subtitle = "more info here") {
  # 动态匹配统计值列(以_est结尾)和使用率列(以_use结尾)
  est_cols <- grep("_est$", names(df_here), value = TRUE)
  use_cols <- grep("_use$", names(df_here), value = TRUE)
  
  df_here %>%
    gt(rowname_col = "age") %>%
    tab_stubhead(label = md("**Statewide**")) %>%
    tab_header(
      title = md(table_title),
      subtitle = table_subtitle
    ) %>%
    fmt_percent(
      columns = all_of(use_cols),
      decimals = 1,
      scale_values = FALSE  # 数据为小数时设为FALSE,整数形式则改为TRUE
    ) %>%
    fmt_number(
      columns = all_of(est_cols),
      decimals = 0
    ) %>%
    tab_spanner_delim(
      delim = "_",
      columns = -"age"  # 自动排除行名列,处理其余所有列
    ) %>%
    return()  # 显式返回gt表格对象
}

关键优化点

  • 显式返回结果:通过return()确保函数输出格式化后的gt对象
  • 动态列匹配:用grep()根据列名后缀自动识别目标列,避免硬编码列号的适配问题
  • 参数化标题:新增标题参数,支持为每个表格自定义标题内容
  • 灵活列选择:用-"age"排除行名列,自动对其余列应用分隔符表头,无需指定固定列范围

批量处理示例

若多个数据框存储在列表中,可通过lapply()批量应用格式化:

# 模拟多数据框场景
df_list <- list(
  example,
  example %>% mutate(across(c(alc_est, marj_est), ~ .x * 1.2)),
  example %>% mutate(across(c(bing_alc, mj_use), ~ .x * 0.8))
)

# 批量格式化所有数据框
formatted_tables <- lapply(df_list, format_use_est, 
                          table_title = "**Substance Use Estimates**",
                          table_subtitle = "Statewide Age Group Breakdown")

# 查看第一个格式化后的表格
formatted_tables[[1]]

注意事项

  • 确保所有待格式化的数据框都包含age列,作为表格行名列
  • 列名需保持xxx_est(统计值)和xxx_use(使用率)的命名规则,否则动态列匹配会失效
  • 如果使用率数据是整数形式(如4.77而非0.0477),需将fmt_percent()中的scale_values参数设为TRUE

内容的提问来源于stack exchange,提问作者KLenny

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 08:06:20