You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用anesrake包计算样本权重报错:x + weights遇非数值参数

解决anesrake报错:Error in x + weights : non-numeric argument to binary operator

问题根源

你当前构建population数据框的方式错误,将两个独立的目标变量(政党认同、意识形态)的类别做了一一绑定,导致weights::wpct生成的目标对象结构不符合anesrake的要求。anesrake需要的是每个变量的边际分布目标,而非变量间的联合分布。

另外,代码中通过dplyr::data_frame构建的population会让ideol和pid的类别按行配对,这不是anesrake期望的输入格式。

修正步骤

  1. 分别构建单个变量的目标分布
    不需要合并两个变量到同一个数据框,直接为每个变量单独创建权重目标,确保类别顺序与数据中的因子类别完全一致。

  2. 确保数据中的变量为因子且类别匹配
    确认df中的pid和ideol是因子类型,且类别标签与目标分布中的完全对应(大小写、拼写一致)。

  3. 简化target列表的创建
    直接针对每个变量的类别和对应比例生成wpct对象,无需借助中间数据框。

修正后的代码

# recoding/labeling target variables
df <- df %>% 
  mutate(pid = case_when(
           partyid7_ %in% c(1,2) ~ "Democrat",
           partyid7_ %in% c(3,4,5) ~ "Independent",
           partyid7_ %in% c(6,7) ~ "Republican"
         ),
         ideol = case_when(
           ideology %in% c(1,2) ~ "Liberal",
           ideology == 3 ~ "Moderate",
           ideology %in% c(4,5) ~ "Conservative"
         )) %>%
  # 转换为因子,确保类别顺序与目标一致
  mutate(pid = factor(pid, levels = c("Democrat", "Independent", "Republican")),
         ideol = factor(ideol, levels = c("Liberal", "Moderate", "Conservative")))

# 检查并移除NA(确保无缺失)
df <- df %>% filter(!is.na(pid), !is.na(ideol))

# 设置单个变量的目标参数
# 意识形态目标:类别+比例
ideo_target <- weights::wpct(
  factor(c("Liberal", "Moderate", "Conservative"), 
         levels = c("Liberal", "Moderate", "Conservative")),
  c(0.25, 0.37, 0.36)
)

# 政党认同目标:类别+比例
pid_target <- weights::wpct(
  factor(c("Democrat", "Independent", "Republican"), 
         levels = c("Democrat", "Independent", "Republican")),
  c(0.28, 0.41, 0.28)
)

# 创建target列表
target <- list(ideol = ideo_target, pid = pid_target)

# 添加唯一caseid(确保为数值型)
df$caseid <- seq_len(nrow(df))

# 运行anesrake
dfrake <- anesrake(
  target,                     
  df,                          
  caseid = df$caseid,               
  cap = 3,                     
  choosemethod = "total",      
  type = "pctlim",             
  pctlim = 0.05                 
)

额外排查点

  • 确认df中pid和ideol的因子类别没有额外的水平(比如拼写错误、大小写不一致),可以用levels(df$pid)和levels(df$ideol)检查。
  • 确保caseid是唯一的数值型向量,没有重复值。
  • 如果仍报错,尝试简化代码:先只用一个变量(比如只加权pid)测试,确认单个变量能运行后再加第二个变量,逐步排查。

内容的提问来源于stack exchange,提问作者sce415

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 02:05:28