使用tbl_svysummary处理复现权重调查对象报错求助
tbl_svysummary处理复现权重survey对象分组报错排查与解决
问题概述
使用tbl_svysummary处理带BRR复现权重的survey对象时,设置by = ambulance分组参数后出现命名不匹配错误,核心原因是内置统计量p.std.error在分组场景下的输出结构与函数预期不兼容,而非tbl_svysummary不支持复现权重。
原代码与报错信息
原代码
library(survey) library(gtsummary) data(scd) repweights<-2*cbind(c(1,0,1,0,1,0), c(1,0,0,1,0,1), c(0,1,1,0,0,1), c(0,1,0,1,1,0)) scdrep<-svrepdesign(data=scd, type="BRR", repweights=repweights, combined.weights=FALSE) scd_svy <- srvyr::as_survey(scdrep) scd_svy |> tbl_svysummary(by = ambulance, include = ESA, statistic = all_categorical() ~ "{p} ({p.std.error})")
报错信息
Error in
mutate():
ℹ In argument:df_stats = pmap(...).
Caused by error inpmap():
ℹ In index: 1.
Caused by error inset_names():
! The size ofnm(5) must be compatible with the size ofx(4).
Runrlang::last_trace()to see where the error occurred.
解决方案
自定义统计函数适配分组场景下的比例及标准误计算,替代内置的p.std.error统计符:
修改后代码
library(survey) library(gtsummary) library(srvyr) library(glue) library(scales) # 自定义统计函数:按分组计算分类变量的加权比例及标准误 calc_prop_se <- function(data, var, by_var) { data |> group_by({{ by_var }}) |> # 对分类变量的每个水平计算比例和标准误 summarize( across({{ var }}, ~{ survey_res <- survey_mean(.x == levels(.x)[1], proportion = TRUE, se = TRUE) glue("{percent(survey_res[[1]], accuracy = 0.1)} ({number(survey_res[[2]], accuracy = 0.01)})") }) ) |> pull({{ var }}) } data(scd) repweights<-2*cbind(c(1,0,1,0,1,0), c(1,0,0,1,0,1), c(0,1,1,0,0,1), c(0,1,0,1,1,0)) scdrep<-svrepdesign(data=scd, type="BRR", repweights=repweights, combined.weights=FALSE) scd_svy <- as_survey(scdrep) scd_svy |> tbl_svysummary( by = ambulance, include = ESA, statistic = all_categorical() ~ "{calc_prop_se(., ESA, ambulance)}" )
原因说明
tbl_svysummary完全支持复现权重的survey对象,报错并非兼容性问题。- 内置的
p.std.error统计量是针对整体样本计算标准误,当使用by参数分组后,该统计量返回的结果数量与函数为分组列分配的命名数量不匹配,从而触发set_names的大小兼容错误。 - 自定义统计函数通过
group_by按分组变量拆分数据,分别计算每组内的比例和标准误,输出结构完全适配tbl_svysummary的分组场景要求。
内容的提问来源于stack exchange,提问作者piggyspiggy
相关产品推荐
相关产品推荐

