You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效生成1亿个指定均值和标准差的正态分布数值?

问题:高效生成大规模指定参数的正态分布数值

需要生成均值为5.357、标准差为2.37的大量正态分布数值,当规模达到10,000,000量级时,random.normalvariate、random.gauss以及循环调用np.random.normal的方式耗时过长,寻求低耗时的最优实现方案。


已测试方法及小规模(1000条数据)结果

1. random.normalvariate

import random
import numpy as np

new_list_normalvariate = [random.normalvariate(5.357, 2.37) for x in range(1000)]
print(new_list_normalvariate[0:10])
print('mean = ', np.mean(new_list_normalvariate))
print('std = ', np.std(new_list_normalvariate))

输出:

[6.576049386450241, 8.62262371117091, 4.921246966899101, 6.751587914411607, 5.6042223736139105, 4.493753810671122, 7.868066836581562, 6.299169672752275, 6.081202725113191, 7.27255885543875]
mean =  5.3337034248054875
std =  2.4124820216611336

2. random.gauss

new_list_gauss = [random.gauss(5.357, 2.37) for x in range(1000)]
print(new_list_gauss[0:10])
print('mean = ', np.mean(new_list_gauss))
print('std = ', np.std(new_list_gauss))

输出:

[4.160280814524453, 8.376767324676795, 8.476968737124544, 6.050223384914485, 2.6635671201126785, 2.4441297408189167, 7.624650437282289, 7.5957096799039485, 1.990806588702878, 1.7821756994741982]
mean =  5.347638951117946
std =  2.374617608342891

3. 循环调用np.random.normal

new_list_np_normal = [np.random.normal(5.357, 2.37) for x in range(1000)]
print(new_list_np_normal[0:10])
print('mean = ', np.mean(new_list_np_normal))
print('std = ', np.std(new_list_np_normal))

输出:

[4.294445875786478, 4.930900785615266, 8.244969311017886, 3.380908919026986, 3.636133194752361, 6.191836517294145, 5.17400630491519, 3.16529157634111, 1.9176117359394778, 8.269659173531764]
mean =  5.417575775284877
std =  2.373787525312793

最优解决方案:NumPy向量化生成

Python循环的解释型开销是大规模生成耗时的核心原因,NumPy的np.random.normal支持直接指定生成数量,底层通过C级向量化运算实现,效率远高于循环调用。

基础实现代码

import numpy as np

# 直接生成10000000个符合要求的正态分布数值
new_list = np.random.normal(loc=5.357, scale=2.37, size=10000000)

# 验证均值与标准差
print('mean = ', np.mean(new_list))
print('std = ', np.std(new_list))

性能升级方案(NumPy 1.17+推荐)

使用np.random.Generator新接口,随机数生成效率更高,且线程安全:

rng = np.random.default_rng()
new_list = rng.normal(loc=5.357, scale=2.37, size=10000000)

额外优化建议

  • 若无需高精度浮点数,指定dtype=np.float32可减少内存占用并进一步提升速度:
    new_list = rng.normal(loc=5.357, scale=2.37, size=10000000, dtype=np.float32)
    
  • 生成后的数据直接以NumPy数组存储,后续统计、计算可直接复用数组接口,避免列表转数组的额外开销。

内容的提问来源于stack exchange,提问作者Khaled DELLAL

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 05:05:22