使用tbl_strata+tbl_summary分层表格异常:统计方式不符预期
问题与解决:tbl_strata分层时连续变量统计方式异常
问题描述
需要生成按group_category和Sex双重分层的表格,要求所有变量以mean(se)统计。但使用tbl_strata结合tbl_summary的分层代码时,delta_IL1B变量却被以n(%)统计,而仅按group_category单分组的代码中该变量统计正常。
单分组正常代码
cyto_table <- ccf.data.merge %>% #data frame select(group_category, delta_TNFa, delta_IFNy, delta_IL1B, delta_IL2, delta_IL6, delta_IL8) %>% #variables I want for my table tbl_summary( by = group_category, #group_category = treatment A and B statistic = list(all_continuous() ~ "{mean} ({se})"), missing = "no") %>% modify_header(label ~ "**Variable**") %>% bold_labels() cyto_table
分层异常代码
cyto_table <- ccf.data.merge %>% select(group_category,Sex,delta_TNFa, delta_IFNy, delta_IL1B, delta_IL2, delta_IL6, delta_IL8) %>% #variables I want for my table dplyr::mutate( Sex = factor(Sex, labels = c("Female", "Male"))) %>% #I am not sure the reasoning behind this line. the original code had this, I'm not gonna argue. tbl_strata( strata = group_category, #here the code is layering the able to be stratified by the treatments (A and B). ~.x %>% #I believe this line is making a function for tbl_summary to pass through. tbl_summary( by = Sex, #finally Sex is added here so that the data is then split by Sex within each treatment. statistic = list(all_continuous() ~ "{mean} ({se})"), missing = "no") %>% modify_header(label ~ "**Variable**") %>% bold_labels()) cyto_table
异常原因
核心原因是分层后的子数据集里,delta_IL1B被tbl_summary自动识别为分类变量:
- 当按
group_category分层后,某个子组内的delta_IL1B可能只有唯一值(比如全部为0),或者数值数量过少,触发了tbl_summary的自动类型判断逻辑,将其归类为分类变量。 - 单分组时,整个数据集的
delta_IL1B是连续分布的,所以被正确识别为连续变量,使用mean(se)统计。
解决方法
1. 强制指定变量类型
在tbl_summary中通过type参数强制将delta_IL1B标记为连续变量,忽略自动识别结果:
cyto_table <- ccf.data.merge %>% select(group_category,Sex,delta_TNFa, delta_IFNy, delta_IL1B, delta_IL2, delta_IL6, delta_IL8) %>% dplyr::mutate(Sex = factor(Sex, labels = c("Female", "Male"))) %>% tbl_strata( strata = group_category, ~.x %>% tbl_summary( by = Sex, # 强制指定delta_IL1B为连续变量 type = list(delta_IL1B ~ "continuous"), statistic = list(all_continuous() ~ "{mean} ({se})"), missing = "no") %>% modify_header(label ~ "**Variable**") %>% bold_labels()) cyto_table
2. 检查分层子数据集的变量分布
可以临时在tbl_strata中添加打印代码,查看每个分层组内delta_IL1B的数值情况,确认是否存在唯一值或类型异常:
cyto_table <- ccf.data.merge %>% select(group_category,Sex,delta_TNFa, delta_IFNy, delta_IL1B, delta_IL2, delta_IL6, delta_IL8) %>% dplyr::mutate(Sex = factor(Sex, labels = c("Female", "Male"))) %>% tbl_strata( strata = group_category, ~{ # 打印当前分层组的delta_IL1B统计信息 cat("当前分层组:", unique(.x$group_category), "\n") print(summary(.x$delta_IL1B)) .x %>% tbl_summary( by = Sex, statistic = list(all_continuous() ~ "{mean} ({se})"), missing = "no") %>% modify_header(label ~ "**Variable**") %>% bold_labels() })
根据打印结果,如果发现子组内delta_IL1B确实无变异,可以考虑合并组或者在统计时注明情况。
3. 提前统一变量类型
在数据预处理阶段,明确将delta_IL1B转换为数值型,避免分层时出现类型转换异常:
cyto_table <- ccf.data.merge %>% select(group_category,Sex,delta_TNFa, delta_IFNy, delta_IL1B, delta_IL2, delta_IL6, delta_IL8) %>% dplyr::mutate( Sex = factor(Sex, labels = c("Female", "Male")), # 强制转换为数值型 delta_IL1B = as.numeric(delta_IL1B) ) %>% tbl_strata( strata = group_category, ~.x %>% tbl_summary( by = Sex, statistic = list(all_continuous() ~ "{mean} ({se})"), missing = "no") %>% modify_header(label ~ "**Variable**") %>% bold_labels()) cyto_table
内容的提问来源于stack exchange,提问作者Tara Mahmood
相关产品推荐
相关产品推荐

