You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于matchit()输出中distance的技术问询:含义及计算对象咨询

Understanding matchit() Output: The distance Metric

Great questions—let’s break this down clearly, since the distance metric in matchit() is a common point of confusion!

1. What exactly does the distance in matchit() output refer to?

In most standard matching workflows (using the default settings of matchit()), the distance values are propensity scores. A propensity score is the estimated probability that a given observation (unit) is assigned to the treatment group, based on its observed covariates (like age, income, or demographic factors).

For example, if you run:

m.out <- matchit(treat ~ age + gender + education, data = my_dataset)

The distance column in m.out$distance will be each observation’s calculated probability of being in the treatment group, derived from the logistic regression model fit to the covariates age, gender, and education.

That said, distance isn’t limited to propensity scores. If you specify a different metric via the distance argument (e.g., distance = "mahalanobis"), or supply a custom distance calculation, the distance values will correspond to that specific metric. But propensity scores are the default and most widely used case.

2. Between which two objects is the distance in matchit() output calculated?

Let’s clarify a common misconception first: The distance values in the matchit() output are individual unit-level scores, not a direct pairwise distance between two units. However, these scores are the building blocks for calculating pairwise distances during the matching process.

Here’s how it works for the default propensity score matching:

  • Each observation gets its own distance (propensity score) value.
  • When the algorithm performs matching (e.g., nearest-neighbor matching), it computes the distance between each treatment unit and every potential control unit using their propensity scores. The most common way to do this is taking the absolute difference between the treatment unit’s score and a control unit’s score. The control unit with the smallest difference becomes the match for the treatment unit.

For other distance metrics like Mahalanobis distance:

  • The matchit() function doesn’t store per-unit distance values in the output. Instead, it directly calculates pairwise Mahalanobis distances between treatment and control units using their raw covariate values (this metric measures how far apart two units are in the multivariate covariate space), and uses those pairwise distances to find matches.

内容的提问来源于stack exchange,提问作者OverFlow Police

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:20:19