离散随机变量转连续:正态近似连续性校正疑问
Hey there! Let's break down exactly why that 2.59 value shows up in your continuity correction step. Here's a step-by-step breakdown tailored to your scenario:
1. First, clarify the discrete nature of your sample mean
Your random variable ( H ) takes discrete values: 1, 2, 3, 4. When you take 50 observations, the total sum ( S = H_1 + H_2 + ... + H_{50} ) is an integer (since each ( H ) is integer). Your sample mean ( \bar{X} = S/50 ), so it can only take values that are multiples of ( 1/50 = 0.02 ) (e.g., 1.00, 1.02, 1.04, ..., 2.50, ..., 3.98, 4.00).
This means ( \bar{X} ) isn't a continuous variable—it jumps in 0.02 increments. The normal distribution we're using for approximation (( N(2.5, 0.025) )) is continuous, so we need to "map" these discrete jumps to continuous intervals.
2. The core of continuity correction for your sample mean
For any discrete variable with step size ( \Delta ), each discrete value ( x ) corresponds to a continuous interval ( [x - \Delta/2, x + \Delta/2) ). In your case:
- Step size ( \Delta = 0.02 )
- Half-step size ( \Delta/2 = 0.01 )
So:
- ( \bar{X} = 2.58 ) maps to ( [2.57, 2.59) )
- ( \bar{X} = 2.60 ) maps to ( [2.59, 2.61) )
3. Why 2.59 is the corrected threshold
Let's say you're calculating a probability like ( P(\bar{X} > 2.59) ) (a common scenario that would lead to this correction). Since ( \bar{X} ) can only take multiples of 0.02, the smallest discrete value greater than 2.59 is 2.60.
All discrete values ( \bar{X} \geq 2.60 ) correspond to the continuous interval union ( [2.59, \infty) ). So when using the normal approximation, we calculate ( P(N(2.5, 0.025) \geq 2.59) )—that's where the 2.59 comes from.
Alternatively, if you were calculating ( P(\bar{X} \leq 2.59) ), the largest discrete value ≤2.59 is 2.58, which maps to ( [2.57, 2.59) ). The union of all such intervals is ( (-\infty, 2.59) ), so we use ( P(N(2.5, 0.025) \leq 2.59) ) for the approximation.
Quick recap
The 2.59 isn't an arbitrary number—it's the boundary of the continuous interval that corresponds to the discrete sample mean values you care about. It comes from adjusting the nearest discrete threshold by half the step size (0.01) of your sample mean.
内容的提问来源于stack exchange,提问作者burn_burn_55

