You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何生成指定均值、方差且符合卡方分布的薪资类模拟数据?

Generating Salary Data with Custom Mean/Variance Using Chi-Square Distribution

Got it, let's break this down. The core issue here is that numpy's np.random.chisquare only lets you specify degrees of freedom (df), but chi-square distributions have fixed relationships between their mean, variance, and df:

  • For a chi-square variable (X \sim \chi^2(k)) (where (k) is df), the mean (E[X] = k)
  • The variance (Var(X) = 2k)

So native chi-square distributions have variance exactly twice the mean. To get a chi-square-derived distribution with your custom mean ((\mu)) and variance ((\sigma^2)), you'll need to apply a linear transformation to the raw chi-square data. Here's how to make it work:

Step 1: Derive the Transformation Parameters

We’ll create a new variable (Y = aX + b), where (X) is the raw chi-square data, and (a)/(b) are scaling/shift parameters to hit your target mean and variance.

From the properties of expectation and variance:

  1. (E[Y] = a \cdot E[X] + b = a \cdot k + b = \mu)
  2. (Var(Y) = a^2 \cdot Var(X) = a^2 \cdot 2k = \sigma^2)

You can choose any positive (k) (df) depending on the skewness you want for your salary data:

  • Smaller (k) = more right-skewed (great for salary data, where a few high earners pull the mean up)
  • Larger (k) = closer to a normal distribution (via the Central Limit Theorem)

A common choice is to set (k = \mu) (matches the target mean to the raw chi-square mean), which simplifies the calculations:

  • (a = \sqrt{\frac{\sigma^2}{2k}} = \sqrt{\frac{\sigma^2}{2\mu}})
  • (b = \mu - a \cdot k = \mu - a \cdot \mu = \mu(1 - a))

Step 2: Implement in Python

Here’s a reusable function that generates your desired salary data, with a safeguard to ensure no negative salaries (since pay can’t be negative):

import numpy as np

def generate_chi_salary(target_mean, target_var, sample_size, df=None):
    # Set df to target_mean if not specified (balances skewness and simplicity)
    k = df if df is not None else target_mean
    
    # Calculate scaling and shift parameters
    scale = np.sqrt(target_var / (2 * k))
    shift = target_mean - (scale * k)
    
    # Generate raw chi-square data
    raw_chisq = np.random.chisquare(df=k, size=sample_size)
    
    # Apply transformation and clamp to non-negative values
    salaries = scale * raw_chisq + shift
    salaries = np.maximum(salaries, 0)
    
    return salaries

Step 3: Test the Function

Let’s verify it works with two scenarios:

Scenario 1: Variance is twice the mean (matches native chi-square)

mu = 5000
var = 10000  # 2 * mu
salaries = generate_chi_salary(mu, var, 10000)

print(f"Mean: {np.mean(salaries):.2f}")    # Should be ~5000
print(f"Variance: {np.var(salaries):.2f}") # Should be ~10000

Scenario 2: Custom variance (not twice the mean)

mu = 6000
var = 18000  # Not 2*mu
salaries = generate_chi_salary(mu, var, 10000)

print(f"Mean: {np.mean(salaries):.2f}")    # Should be ~6000
print(f"Variance: {np.var(salaries):.2f}") # Should be ~18000

Key Notes

  • Skewness Control: If you want more pronounced right skew (like real-world salaries where top earners are outliers), set a smaller df (e.g., df=10). For more normal-looking data, use a larger df.
  • Non-Negative Check: The linear transformation might produce negative values (especially with small df), so np.maximum(salaries, 0) ensures your salary data makes sense.

内容的提问来源于stack exchange,提问作者timmy-o-toole

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 19:07:37