You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

scipy.stats.expon.rvs()与numpy.random.exponential()的差异及绘图横轴不同原因

无安打比赛的出现频率统计

美国职业棒球大联盟现代时期(1901-2015年)每场无安打比赛之间的举办场次间隔,存储在数组nohitter_times中。
如果假设无安打比赛符合泊松过程,那么两场无安打比赛的间隔服从指数分布。指数分布仅有单个参数,记作$τ$,即典型间隔时间。
能让指数分布最优拟合数据的参数$τ$取值,就是无安打比赛的平均间隔时间(时间单位为比赛场次数)。

# Here you go with the data
nohitter_times = np.array([ 843, 1613, 1101,  215,  684,  814,  278,  324,  161,  219,  545,
        715,  966,  624,   29,  450,  107,   20,   91, 1325,  124, 1468,
        104, 1309,  429,   62, 1878, 1104,  123,  251,   93,  188,  983,
        166,   96,  702,   23,  524,   26,  299,   59,   39,   12,    2,
        308, 1114,  813,  887,  645, 2088,   42, 2090,   11,  886, 1665,
       1084, 2900, 2432,  750, 4021, 1070, 1765, 1322,   26,  548, 1525,
         77, 2181, 2752,  127, 2147,  211,   41, 1575,  151,  479,  697,
        557, 2267,  542,  392,   73,  603,  233,  255,  528,  397, 1529,
       1023, 1194,  462,  583,   37,  943,  996,  480, 1497,  717,  224,
        219, 1531,  498,   44,  288,  267,  600,   52,  269, 1086,  386,
        176, 2199,  216,   54,  675, 1243,  463,  650,  171,  327,  110,
        774,  509,    8,  197,  136,   12, 1124,   64,  380,  811,  232,
        192,  731,  715,  226,  605,  539, 1491,  323,  240,  179,  702,
        156,   82, 1397,  354,  778,  603, 1001,  385,  986,  203,  149,
        576,  445,  180, 1403,  252,  675, 1351, 2983, 1568,   45,  899,
       3260, 1025,   31,  100, 2055, 4043,   79,  238, 3931, 2351,  595,
        110,  215,    0,  563,  206,  660,  242,  577,  179,  157,  192,
        192, 1848,  792, 1693,   55,  388,  225, 1134, 1172, 1555,   31,
       1582, 1044,  378, 1687, 2915,  280,  765, 2819,  511, 1521,  745,
       2491,  580, 2072, 6450,  578,  745, 1075, 1103, 1549, 1520,  138,
       1202,  296,  277,  351,  391,  950,  459,   62, 1056, 1128,  139,
        420,   87,   71,  814,  603, 1349,  162, 1027,  783,  326,  101,
        876,  381,  905,  156,  419,  239,  119,  129,  467])

第一种实现方法(scipy)

import scipy.stats as stats

# computing the distribution parameter
avg_interval = np.mean(nohitter_times)

# Set the seed
np.random.seed(42)
# Simulating the distribution
rvs = stats.expon.rvs(avg_interval, size=100000)

#Plotting the distribution
#sns.histplot(rvs, kde=True, bins=100, color='skyblue', stat='density');
_ = plt.hist(rvs, bins=50, density=True, histtype="step")
_ = plt.xlabel('Games between no-hitters')
_ = plt.ylabel('PDF');

scipy方法绘制结果

第二种实现方法(numpy)

# Seed random number generator
np.random.seed(42)

# Compute mean no-hitter time: tau
tau = np.mean(nohitter_times)

# Draw out of an exponential distribution with parameter tau: inter_nohitter_time
inter_nohitter_time = np.random.exponential(tau, 100000)

# Plot the PDF and label axes
_ = plt.hist(inter_nohitter_time, bins=50, density=True, histtype="step")
_ = plt.xlabel('Games between no-hitters')
_ = plt.ylabel('PDF')

numpy方法绘制结果

差异原因

两个库指数分布函数的参数定义不同是核心原因:

  • numpy.random.exponential的第一个入参就是指数分布的尺度参数scale,也就是分布均值,你传入tau的用法正确,生成的样本均值为tau,取值范围从0到数千,符合实际数据特征。
  • scipy.stats.expon.rvs的第一个入参是位置参数loc,第二个入参才是代表均值的尺度参数scale。你只传了一个参数的情况下,默认loc=avg_interval、scale=1,生成的样本是loc + 标准指数分布(均值为1),所有样本都集中在avg_interval(约735)附近,取值范围仅在735~745之间,自然和numpy结果的x轴完全不同。

如果要让scipy生成的样本和numpy一致,修改调用代码即可:

# 显式指定scale参数
rvs = stats.expon.rvs(scale=avg_interval, size=100000)
# 或者指定loc为0,第二个参数传scale
rvs = stats.expon.rvs(0, avg_interval, size=100000)

内容的提问来源于stack exchange,提问作者Osama Hamdy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 13:18:01