You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R中分数probit模型聚类标准误计算异常问题排查

Fixing Clustered Standard Errors for Quasibinomial Probit Models in R

It’s common to run into issues when rolling your own clustered SE functions—small oversights in the sandwich estimator math can throw off results entirely. Let’s break down why your custom function isn’t working and how to fix it properly.

The Problem with Your Custom Function

Your approach is on the right track, but there are a few potential pitfalls:

  • Incorrect combination of bread and meat: While you’re calculating cluster-level score sums (uj), relying on sandwich(fit, meat = ...) might not correctly integrate the bread matrix (inverse of expected Fisher information) with your custom meat for quasibinomial models, especially with a probit link.
  • Dispersion parameter handling: Quasibinomial models estimate a dispersion parameter, which your custom function might not be accounting for properly when combining components.

The good news is you don’t need to reinvent the wheel—there’s a well-tested function in the sandwich package specifically for clustered SEs: vcovCL.

Corrected Approach Using vcovCL

Replace your custom clrobustse function with this simplified version that leverages the built-in vcovCL function, which handles quasibinomial models, probit links, and dispersion automatically:

require(sandwich)
require(lmtest)

clrobustse <- function(fit, cl){
  # Use sandwich's built-in clustered variance estimator
  vcov_cluster <- vcovCL(fit, cluster = cl)
  # Return coefficient tests with clustered SEs
  coeftest(fit, vcov = vcov_cluster)
}

Then run your model exactly as before:

fit <- glm(dep ~ exp1 + exp2 + exp3, data = df, fam = quasibinomial("probit"))
clrobustse(fit, df$cluster)

This should produce results that match Stata’s clustered SE output.

Why This Works

The vcovCL function:

  1. Correctly computes cluster-level score contributions, accounting for the probit link’s derivative and quasibinomial dispersion.
  2. Applies the appropriate degrees-of-freedom correction ((M/(M-1))*((N-1)/(N-K))) automatically.
  3. Properly combines the bread and meat components of the sandwich estimator to produce valid clustered standard errors.

If you’re curious about the math behind it, vcovCL follows the standard clustered variance formula:
[ V_{\text{clustered}} = \frac{M}{M-1} \cdot \frac{N-1}{N-K} \cdot (X'WX)^{-1} \left( \sum_{j=1}^M u_j u_j' \right) (X'WX)^{-1} ]
where (u_j) is the sum of score vectors for cluster (j), and (W) is the weight matrix specific to your quasibinomial probit model.

Content of the question来源于stack exchange,提问作者Antti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:34:02