定义Python函数时如何向random.choice传递随机数数组及内存报错解决
报错根因
你传入np.random.choice的第二个参数n是一维数组类型(N是np.random.poisson生成的包含多个泊松随机值的数组),而np.random.choice的size参数如果接收数组,会被解析为输出数组的多维维度,而不是为数组内每个值单独生成对应长度的一维序列。你报错信息里的9维数组就是传入了长度为9的N数组导致的,总元素数为所有维度值的乘积,自然会出现PiB级的内存申请报错。
解决方法
有两种常用实现方式,都可以避免内存错误:
方法1:遍历每个N值单独计算(可读性高,适合N样本量不大的场景)
import numpy as np # 定义随机停和计算函数 def tn_fun(n): return sum(np.random.choice([50, 100, 200], n, replace=True, p=[0.3, 0.5, 0.2])) # 生成10000个lambda=50的泊松随机值 N = np.random.poisson(50, 10000) # 遍历每个N值计算对应随机停和 TN = [tn_fun(n) for n in N] print('Sample mean of the randomly stopped sum TN is',np.mean(TN)) print('Sample variance of the randomly stopped sum TN is', np.var(TN))
方法2:向量化实现(效率更高,适合大样本量场景)
利用numpy广播特性一次性生成所有需要的随机数,避免循环开销:
import numpy as np # 生成10000个lambda=50的泊松随机值 N = np.random.poisson(50, 10000) max_n = N.max() # 生成维度为[样本数, 最大N值]的所有随机采样值 all_vals = np.random.choice([50, 100, 200], size=(len(N), max_n), replace=True, p=[0.3, 0.5, 0.2]) # 生成掩码只保留每个样本前N个有效取值,求和得到随机停和 mask = np.arange(max_n) < N[:, None] TN = (all_vals * mask).sum(axis=1) print('Sample mean of the randomly stopped sum TN is',np.mean(TN)) print('Sample variance of the randomly stopped sum TN is', np.var(TN))
内容的提问来源于stack exchange,提问作者Lakshman
相关产品推荐
相关产品推荐

