有序样本CDF修正估计式推导问询:(i-3/8)/(n+1/4)来源解析
Great question! That adjusted CDF estimate you mentioned — $P(X\leq x_{(i)})\approx\frac{i-\frac{3}{8}}{n+\frac{1}{4}}$ — is known as Blom's Formula, a widely used correction for empirical CDF/quantile estimates, especially for uniformly distributed data. Let’s break down how it’s derived and why it outperforms the simpler $\frac{i}{n}$ approach.
The Problem with the Basic $\frac{i}{n}$ Estimate
First, let’s recall why the naive estimate isn’t ideal. For a sample of size $n$ sorted into $x_{(1)}\leq x_{(2)}\leq\ldots\leq x_{(n)}$ from a uniform $U(0,1)$ distribution, the expected value of the $i$-th order statistic $X_{(i)}$ is $\mathbb{E}[X_{(i)}] = \frac{i}{n+1}$, not $\frac{i}{n}$. The $\frac{i}{n}$ estimate introduces a systematic bias, especially in small samples. For example, with $n=5$ and $i=1$, $\frac{1}{5}=0.2$, but the true expected value is $\frac{1}{6}\approx0.1667$ — a noticeable gap.
Deriving Blom's Formula
Blom’s goal was to find a corrected estimate of the form $\frac{i - a}{n + b}$ that minimizes the mean squared error (MSE) between the estimate and the true expected value of the order statistic. Here’s the step-by-step reasoning:
- Start with the general form of the corrected estimate: $\hat{F}(x_{(i)}) = \frac{i - a}{n + b}$
- We want this estimate to closely match the true expectation $\mathbb{E}[X_{(i)}] = \frac{i}{n+1}$, while also accounting for the variance of the order statistic ($\text{Var}(X_{(i)}) = \frac{i(n-i)}{(n+1)^2(n+2)}$)
- By solving for the values of $a$ and $b$ that minimize the MSE (balancing bias reduction and variance control), Blom derived the optimal constants:
- $a = \frac{3}{8}$
- $b = \frac{1}{4}$
An alternative way to frame this is using Edgeworth expansions — a method to approximate the distribution of order statistics beyond the first asymptotic term. When expanding the distribution of $X_{(i)}$ for large $n$, the leading bias term can be eliminated by choosing these exact constants, leading to the formula you’re asking about.
Why This Correction is "Better"
Unlike the naive $\frac{i}{n}$ estimate, Blom’s formula:
- Reduces small-sample bias significantly, aligning the estimate much closer to the true expected value of the order statistic
- Maintains a good balance between bias and variance (other corrections like $\frac{i-0.5}{n}$ can overcorrect variance in some cases)
- Works well even when transforming to non-uniform distributions (e.g., estimating quantiles for normal data by first applying the probability integral transform to uniform)
内容的提问来源于stack exchange,提问作者Frank Vel

