You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tbl_svysummary处理复现权重调查对象报错求助

tbl_svysummary处理复现权重survey对象分组报错排查与解决

问题概述

使用tbl_svysummary处理带BRR复现权重的survey对象时,设置by = ambulance分组参数后出现命名不匹配错误,核心原因是内置统计量p.std.error在分组场景下的输出结构与函数预期不兼容,而非tbl_svysummary不支持复现权重。

原代码与报错信息

原代码

library(survey)
library(gtsummary)

data(scd)
repweights<-2*cbind(c(1,0,1,0,1,0), c(1,0,0,1,0,1), c(0,1,1,0,0,1),
                    c(0,1,0,1,1,0))
scdrep<-svrepdesign(data=scd, type="BRR", repweights=repweights, combined.weights=FALSE)

scd_svy <- srvyr::as_survey(scdrep) 

scd_svy |>
  tbl_svysummary(by = ambulance,
                 include = ESA,
                 statistic = all_categorical() ~ "{p} ({p.std.error})")

报错信息

Error in mutate():
ℹ In argument: df_stats = pmap(...).
Caused by error in pmap():
ℹ In index: 1.
Caused by error in set_names():
! The size of nm (5) must be compatible with the size of x (4).
Run rlang::last_trace() to see where the error occurred.

解决方案

自定义统计函数适配分组场景下的比例及标准误计算,替代内置的p.std.error统计符:

修改后代码

library(survey)
library(gtsummary)
library(srvyr)
library(glue)
library(scales)

# 自定义统计函数:按分组计算分类变量的加权比例及标准误
calc_prop_se <- function(data, var, by_var) {
  data |>
    group_by({{ by_var }}) |>
    # 对分类变量的每个水平计算比例和标准误
    summarize(
      across({{ var }}, ~{
        survey_res <- survey_mean(.x == levels(.x)[1], proportion = TRUE, se = TRUE)
        glue("{percent(survey_res[[1]], accuracy = 0.1)} ({number(survey_res[[2]], accuracy = 0.01)})")
      })
    ) |>
    pull({{ var }})
}

data(scd)
repweights<-2*cbind(c(1,0,1,0,1,0), c(1,0,0,1,0,1), c(0,1,1,0,0,1),
                    c(0,1,0,1,1,0))
scdrep<-svrepdesign(data=scd, type="BRR", repweights=repweights, combined.weights=FALSE)

scd_svy <- as_survey(scdrep) 

scd_svy |>
  tbl_svysummary(
    by = ambulance,
    include = ESA,
    statistic = all_categorical() ~ "{calc_prop_se(., ESA, ambulance)}"
  )

原因说明

  • tbl_svysummary完全支持复现权重的survey对象,报错并非兼容性问题。
  • 内置的p.std.error统计量是针对整体样本计算标准误,当使用by参数分组后,该统计量返回的结果数量与函数为分组列分配的命名数量不匹配,从而触发set_names的大小兼容错误。
  • 自定义统计函数通过group_by按分组变量拆分数据,分别计算每组内的比例和标准误,输出结构完全适配tbl_svysummary的分组场景要求。

内容的提问来源于stack exchange,提问作者piggyspiggy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 18:16:25