You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

双重差分法(Difference-in-Differences)与熵匹配结合的样本匹配方法咨询

Entropy Balancing for DID with Imbalanced Covariates (Panel Data Setup)

Great question—entropy balancing is an excellent choice here to address covariate imbalance between your treatment and control groups, especially since pre-treatment differences threaten the parallel trend assumption critical for DID validity. Let’s walk through the step-by-step approach tailored to your 5 pre/5 post-treatment panel data with ~200 firms:

1. First: Curate the Right Covariates for Balancing

Focus exclusively on pre-treatment covariates—this is non-negotiable for avoiding endogeneity. You want to balance variables that:

  • Show significant pre-treatment differences between treatment and control groups (the ones you’ve already identified)
  • Are theoretically linked to both your treatment assignment and outcome variable (e.g., firm size, age, pre-treatment profitability, leverage, industry classification)
  • For panel data, use either the average of pre-treatment values (across the 5 pre-periods) or the most recent pre-treatment value—avoid mixing in post-treatment data at all costs.

2. Implement Entropy Balancing: Core Steps

Entropy balancing works by assigning weights to control group firms such that the weighted control distribution matches the treatment group’s covariate distribution (you can specify mean, variance, or even higher-order moment constraints). Here’s how to execute it:

  • Set balance targets: For DID, start with matching first moments (means) of all key covariates. If you notice significant variance differences between groups, add second-moment constraints to balance spreads too.
  • Calculate weights: Use entropy minimization to find weights that meet your balance targets while keeping the weight distribution as "spread out" as possible (to avoid over-reliance on a small subset of control firms).
    Example R code with the ebalance package:
    library(ebalance)
    # Subset to pre-treatment data only for covariate balancing
    pre_data <- df[df$post == 0, ]
    # Define treatment indicator and pre-treatment covariates
    treat_vec <- pre_data$treat  # 1 = treatment, 0 = control
    cov_matrix <- pre_data[, c("avg_size_pre", "avg_roa_pre", "firm_age", "industry_code")]
    # Run entropy balancing
    eb_output <- ebalance(Treatment = treat_vec, X = cov_matrix, base.weight = rep(1, nrow(pre_data)))
    # Extract weights (match length to full panel data)
    df$entropy_weights <- ifelse(df$treat == 1, 1, eb_output$w[match(df$firm_id, pre_data$firm_id)])
    
  • Validate balance: After calculating weights, verify that all covariates have standardized mean differences (SMD) below 0.1 (a common threshold) between weighted control and treatment groups. You can also run t-tests to confirm no statistically significant differences, or plot density overlays for continuous covariates.

3. Integrate Weights into Your DID Regression

Once you have valid weights, incorporate them into your DID model using weighted least squares (WLS). For panel data, you’ll still want to include firm and time fixed effects to control for unobserved heterogeneity:
Example with fixed effects (using lfe package for efficiency):

library(lfe)
# DID model with entropy weights
did_model <- felm(
  outcome ~ treat * post | firm_id + year | 0 | firm_id,
  data = df,
  weights = df$entropy_weights
)
summary(did_model)

The key coefficient here is the interaction term treat:post—this gives your weighted average treatment effect (WATE) after balancing covariates.

4. Critical Checks for Your Setup

  • Avoid post-treatment covariates: Never include variables measured after treatment starts in your entropy balancing—this would introduce bias by conflating treatment effects with pre-existing differences.
  • Panel vs. cross-sectional matching: Entropy balancing for DID is typically done at the cross-sectional level (matching firms based on their pre-treatment profiles) rather than逐期. This aligns with the parallel trend assumption, which focuses on pre-treatment trajectories.
  • Weight sanity: Check that no control firm has an extremely high weight (e.g., >10) or that weights aren’t concentrated in just a few firms. If this happens, you may need to reduce the number of covariates or relax balance constraints (e.g., drop higher-order moments).

内容的提问来源于stack exchange,提问作者user194011

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:39:21