求助:gtsummary报错age变量statistic参数需为命名函数列表
问题
调用gtsummary::modify_header()时触发错误:
! Error in the argument statistic for variable "age". ℹ Value must be a named list of functions
已知age是数值型连续变量,使用的tbl_summary()代码如下:
gtsummary::tbl_summary( by = variable_x, include = c(age, sex, bmibl), # Variables to summarize type = list(c(bmibl, age) ~ "continuous", sex ~ "categorical" ), statistic = list( age ~ "{mean} ({sd})", # Specify mean and SD for age bmibl ~ "{mean} ({sd})", all_categorical() ~ "{n} ({p}%)" # Counts (percentages) for categorical variables ), digits = list( all_categorical() ~ c(0, 1) # 0 decimal for counts, 1 for percentages ), missing = "no" )
移除statistic列表中的age和bmibl后,会得到非mean(sd)格式的默认统计量,且相同版本的R与gtsummary仅在单个实例上出现该错误。
解决方案
1. 修正statistic参数的语法规范
gtsummary对statistic参数的公式写法有严格要求,单个变量的统计量指定需确保公式解析无歧义。推荐用all_continuous()统一指定连续变量的统计格式,避免单个变量的语法冲突:
gtsummary::tbl_summary( by = variable_x, include = c(age, sex, bmibl), type = list(c(bmibl, age) ~ "continuous", sex ~ "categorical"), statistic = list( all_continuous() ~ "{mean} ({sd})", all_categorical() ~ "{n} ({p}%)" ), digits = list( all_categorical() ~ c(0, 1), all_continuous() ~ c(1, 1) # 可选:为连续变量指定小数位数 ), missing = "no" )
若需单独为某一连续变量指定特殊统计量,需保证公式写法完全正确:
statistic = list( age ~ "{mean} ({sd})", bmibl ~ "{median} ({p25}, {p75})", # 示例:bmibl使用中位数+四分位距 all_categorical() ~ "{n} ({p}%)" )
2. 排查实例环境的隐性冲突
仅单个实例报错的核心原因是环境差异,需逐一排查:
- 检查是否存在同名自定义函数:比如自定义了
mean/sd函数,覆盖了基础包的内置函数,运行exists("mean", mode = "function")验证,若为TRUE则删除该自定义函数 - 检查包依赖冲突:确保实例中加载的
dplyr、tidyselect等包版本与正常实例一致,或暂时卸载其他可能干扰公式解析的包 - 重置工作环境:运行
rm(list = ls()); lapply(names(sessionInfo()$otherPkgs), detach, character.only = TRUE); library(gtsummary)后重新测试
3. 确认变量类型的一致性
验证报错实例中age和bmibl的变量类型确实为numeric:
class(age) class(bmibl)
若为字符型或其他非数值类型,需先转换:
df$age <- as.numeric(df$age) df$bmibl <- as.numeric(df$bmibl)
内容的提问来源于stack exchange,提问作者R-newbie
相关产品推荐
相关产品推荐

