You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言mutate函数报错:propensity_gam维度不匹配问题求助

倾向得分权重计算R代码报错排查与解决

错误原因

报错核心是分组操作导致predict返回结果长度不匹配:

  • 你用group_by(clozapine.or.not)把数据集分成两组,其中第一组有3152行,但predict(mod, type = "response")返回了全数据集的6221个预测值,mutate要求当前组的变量长度必须和组内行数一致,因此触发错误。
  • 更关键的是:倾向得分(propensity score)是针对全样本每个个体计算的(即每个样本被分到处理组的概率),完全不需要分组计算,分组操作本身就是错误的。

解决步骤

1. 移除多余的分组操作

直接在全数据集上计算倾向得分和权重,不需要group_by:

# 修正后的代码:去掉group_by操作
df_all <- df_all %>%
  mutate(propensity_gam = predict(mod, type = "response"),
         weight_gam = 1 / propensity_gam * death_yn + 
           1 / (1 - propensity_gam) * (1 - death_yn)
         )

2. 额外验证(可选但推荐)

  • 确认mod模型是基于df_all训练的,如果模型是用其他子集训练的,需要显式指定newdata = df_all:
    propensity_gam = predict(mod, newdata = df_all, type = "response")
    
  • 检查death_yn是否为0/1的二分类变量,这是权重公式成立的前提。
  • 检查propensity_gam是否存在0或1的极端值(会导致除以0报错),如果有可以用ifelse处理:
    weight_gam = ifelse(death_yn == 1, 
                        1 / pmax(propensity_gam, 1e-6), 
                        1 / pmax(1 - propensity_gam, 1e-6))
    

内容的提问来源于stack exchange,提问作者oohsehun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 20:35:25