tbl_summary函数处理过滤后数据异常问题求助
问题原因与解决方法
你的问题出在Disease_subtype列是因子类型——虽然你过滤掉了取值为0和4的行,但因子原本的水平(0、1、2、3、4)依然保留,tbl_summary会默认展示所有因子水平,没有对应数据的水平就会显示为NA。
解决步骤
直接在过滤后重置因子水平,去掉那些无对应数据的水平即可,以下是两种常用实现方式:
方式1:用基础R的droplevels()
修改后的代码:
table0 <- df.wide.paired.filtered %>% filter(Disease_subtype %in% c(1,2,3)) %>% # 简化过滤条件,和原逻辑等价 mutate(Disease_subtype = droplevels(Disease_subtype)) %>% # 重置因子水平,删除无数据的水平 tbl_summary(include = c(Sex, Age, Disease_subtype, EDSS, MS_Tx, T25FW, HPT, domHPT, nondomHPT, LCVA, SDMT, PASAT), by = Disease_subtype) %>% add_p() %>% bold_labels() table0
方式2:用forcats包的fct_drop()(更简洁)
如果已经加载了tidyverse系列包(包含forcats),可以这样写:
table0 <- df.wide.paired.filtered %>% filter(Disease_subtype %in% c(1,2,3)) %>% mutate(Disease_subtype = fct_drop(Disease_subtype)) %>% tbl_summary(include = c(Sex, Age, Disease_subtype, EDSS, MS_Tx, T25FW, HPT, domHPT, nondomHPT, LCVA, SDMT, PASAT), by = Disease_subtype) %>% add_p() %>% bold_labels() table0
内容的提问来源于stack exchange,提问作者Seon Choi
相关产品推荐
相关产品推荐

