You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用gtsummary的by参数时,如何排除无观测值的列?

问题出在career_stage_s1_v2大概率是因子类型,即便筛选后,原因子的8个水平依然被保留,导致gtsummary生成表格时会显示所有水平列。以下两种方法可解决该问题:

方法1:重置因子水平(推荐,从数据源头处理)

在筛选后重置career_stage_s1_v2的因子水平,仅保留当前数据中存在的取值:

table10 <- select(df, career_stage_s1_v2, mhcsf_emotional_wb, mhcsf_social_wb, mhcsf_psychological_wb) %>%
  filter(career_stage_s1_v2 %in% c("Student", "Currently practicing") & 
           (mhcsf_emotional_wb != "Incomplete" | mhcsf_social_wb != "Incomplete" | mhcsf_psychological_wb != "Incomplete")) %>%
  # 重置因子水平,仅保留有观测的取值
  mutate(career_stage_s1_v2 = factor(career_stage_s1_v2)) %>%
  tbl_summary(
    by = career_stage_s1_v2,
    label = list(career_stage_s1_v2 ~ "Student, Currently practicing"),
    sort = all_categorical() ~ "frequency"
  ) %>%
  modify_caption("**Well Being - emotional, social, or psychological well-being domains by career stage**") %>%
  bold_labels()
table10 

方法2:用gtsummary内置函数过滤列

若不想修改原数据,可在生成表格后,通过modify_column_filter过滤掉观测数为0的列:

table10 <- select(df, career_stage_s1_v2, mhcsf_emotional_wb, mhcsf_social_wb, mhcsf_psychological_wb) %>%
  filter(career_stage_s1_v2 %in% c("Student", "Currently practicing") & 
           (mhcsf_emotional_wb != "Incomplete" | mhcsf_social_wb != "Incomplete" | mhcsf_psychological_wb != "Incomplete")) %>%
  tbl_summary(
    by = career_stage_s1_v2,
    label = list(career_stage_s1_v2 ~ "Student, Currently practicing"),
    sort = all_categorical() ~ "frequency"
  ) %>%
  # 过滤掉观测数为0的分组列
  modify_column_filter(columns = all_stat_cols(), ~ !is.na(.x) & .x > 0) %>%
  modify_caption("**Well Being - emotional, social, or psychological well-being domains by career stage**") %>%
  bold_labels()
table10 

内容的提问来源于stack exchange,提问作者Greg Kawchuk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 10:20:10