如何生成指定均值、方差且符合卡方分布的薪资类模拟数据?
Got it, let's break this down. The core issue here is that numpy's np.random.chisquare only lets you specify degrees of freedom (df), but chi-square distributions have fixed relationships between their mean, variance, and df:
- For a chi-square variable (X \sim \chi^2(k)) (where (k) is df), the mean (E[X] = k)
- The variance (Var(X) = 2k)
So native chi-square distributions have variance exactly twice the mean. To get a chi-square-derived distribution with your custom mean ((\mu)) and variance ((\sigma^2)), you'll need to apply a linear transformation to the raw chi-square data. Here's how to make it work:
Step 1: Derive the Transformation Parameters
We’ll create a new variable (Y = aX + b), where (X) is the raw chi-square data, and (a)/(b) are scaling/shift parameters to hit your target mean and variance.
From the properties of expectation and variance:
- (E[Y] = a \cdot E[X] + b = a \cdot k + b = \mu)
- (Var(Y) = a^2 \cdot Var(X) = a^2 \cdot 2k = \sigma^2)
You can choose any positive (k) (df) depending on the skewness you want for your salary data:
- Smaller (k) = more right-skewed (great for salary data, where a few high earners pull the mean up)
- Larger (k) = closer to a normal distribution (via the Central Limit Theorem)
A common choice is to set (k = \mu) (matches the target mean to the raw chi-square mean), which simplifies the calculations:
- (a = \sqrt{\frac{\sigma^2}{2k}} = \sqrt{\frac{\sigma^2}{2\mu}})
- (b = \mu - a \cdot k = \mu - a \cdot \mu = \mu(1 - a))
Step 2: Implement in Python
Here’s a reusable function that generates your desired salary data, with a safeguard to ensure no negative salaries (since pay can’t be negative):
import numpy as np def generate_chi_salary(target_mean, target_var, sample_size, df=None): # Set df to target_mean if not specified (balances skewness and simplicity) k = df if df is not None else target_mean # Calculate scaling and shift parameters scale = np.sqrt(target_var / (2 * k)) shift = target_mean - (scale * k) # Generate raw chi-square data raw_chisq = np.random.chisquare(df=k, size=sample_size) # Apply transformation and clamp to non-negative values salaries = scale * raw_chisq + shift salaries = np.maximum(salaries, 0) return salaries
Step 3: Test the Function
Let’s verify it works with two scenarios:
Scenario 1: Variance is twice the mean (matches native chi-square)
mu = 5000 var = 10000 # 2 * mu salaries = generate_chi_salary(mu, var, 10000) print(f"Mean: {np.mean(salaries):.2f}") # Should be ~5000 print(f"Variance: {np.var(salaries):.2f}") # Should be ~10000
Scenario 2: Custom variance (not twice the mean)
mu = 6000 var = 18000 # Not 2*mu salaries = generate_chi_salary(mu, var, 10000) print(f"Mean: {np.mean(salaries):.2f}") # Should be ~6000 print(f"Variance: {np.var(salaries):.2f}") # Should be ~18000
Key Notes
- Skewness Control: If you want more pronounced right skew (like real-world salaries where top earners are outliers), set a smaller
df(e.g.,df=10). For more normal-looking data, use a largerdf. - Non-Negative Check: The linear transformation might produce negative values (especially with small
df), sonp.maximum(salaries, 0)ensures your salary data makes sense.
内容的提问来源于stack exchange,提问作者timmy-o-toole

