如何在Matchit企业层面分析中避免PSM匹配出现重复样本?
Got it, let's tackle this problem you're facing with Matchit. The core issue here is your matching rules don't enforce unique firm identifiers, so the same company's data from other years is ending up as controls for your treated firm-year observations. Here's how to lock this down:
1. First, confirm your firm identifier
Make sure your firm_data has a column that uniquely identifies each company—let's call it firm_id (if you don't have one, create it using something like a company name + tax ID that won't repeat across firms).
2. Update your Matchit code to block same-firm matches
You need to add a constraint that prevents treated observations from being matched to any control observations from the same firm. There are two reliable ways to do this:
Method 1: Use a custom pool of control candidates
This approach explicitly limits each treated firm's potential matches to other firms only. We'll use Matchit's pool parameter to define allowed controls for each treated observation:
library(MatchIt) library(dplyr) # For easier data handling psm_out <- matchit(treat ~ firm_variables, data = firm_data, method = "nearest", exact = c("year", "industry"), ratio = 1, # Custom function to restrict controls to other firms pool = function(data, treat, ...) { lapply(which(treat == 1), function(i) { # Grab the firm ID of the current treated observation current_firm <- data$firm_id[i] # Return indices of control firms that aren't the same as current_firm which(treat == 0 & data$firm_id != current_firm) }) })
This tells Matchit: "For every treated firm-year, only consider control observations from completely different companies (same year and industry, per your exact match rules)."
Method 2: Use Mahalanobis distance with a firm ID caliper
If you prefer a more streamlined approach, you can use Mahalanobis distance and set a strict caliper on firm_id to block same-firm matches:
library(MatchIt) psm_out <- matchit(treat ~ firm_variables, data = firm_data, method = "nearest", exact = c("year", "industry"), ratio = 1, distance = "mahalanobis", # Block matches where firm_id is identical (caliper of 0 for firm_id) caliper = list(firm_id = 0))
The caliper here ensures that Matchit won't pair a treated observation with a control that has the same firm_id—since the distance between identical firm IDs is 0, we're saying "no matches allowed at this distance".
3. Verify your matches worked
After running the match, double-check that no same-firm pairs slipped through:
# Extract the matched dataset matched_data <- match.data(psm_out) # Check for any matched pairs where both observations are from the same firm same_firm_check <- matched_data %>% group_by(subclass) %>% summarise(same_firm = n_distinct(firm_id) == 1) %>% filter(same_firm) if (nrow(same_firm_check) == 0) { cat("Success! No same-firm cross-year matches found.") } else { cat("Oops, still have same-firm matches—double-check your firm ID variable and parameters.") }
Quick notes to avoid hiccups
- If you're still seeing different years in matches (like your A2016 vs A2010 example), double-check your
yearvariable's type—it should be numeric or a factor, not a messy character string that Matchit can't match exactly. - Replace
firm_idwith whatever unique firm identifier you use in your data (e.g.,company_id,firm_code). - If you end up with too few control matches, you can try increasing the
ratioparameter (e.g.,ratio=2to get two controls per treated observation) or adjusting yourfirm_variables—just make sure you don't sacrifice the rigor of your analysis.
内容的提问来源于stack exchange,提问作者user12157566

