Python绘制浮点列表概率分布:求各区间概率和为1的绘图工具
问题分析与解决方案
你的核心问题是生成的ans列表是加权重复采样的结果(每个原始val被重复repeats次),直接用普通直方图统计会因为重复计数导致概率/密度偏差,不需要额外工具,用matplotlib或seaborn就能解决。
问题根源
原代码中,每个val被重复repeats次加入ans,相当于每个原始样本的权重是repeats,但普通直方图会把每一次重复都当成独立样本,导致重复次数多的样本被过度计数,最终得到的密度数值偏大。
修正方案
不需要生成庞大的ans列表,直接保存原始val和对应的权重repeats,通过直方图工具的weights参数实现加权统计:
1. 修正数据生成逻辑(避免bug+保存权重)
import math import numpy as np import matplotlib.pyplot as plt import seaborn as sns num = 100000 T = 4.5 * math.pi num_bins = 50 vals = [] weights = [] for i in range(num): val = 0.5 * np.random.chisquare(1) + np.random.exponential(1) q = np.random.randn() # 处理q为负/零的情况,避免生成无效重复次数 if q <= 0: continue repeats = int(T / (2 * q)) if repeats <= 0: continue vals.append(val) weights.append(repeats) # 转换为numpy数组方便后续处理 vals = np.array(vals) weights = np.array(weights)
2. 绘制「区间概率和为1」的直方图
使用seaborn的histplot(推荐替代已弃用的distplot),设置stat='probability'即可让所有区间的概率和为1:
sns.histplot( vals, weights=weights, bins=num_bins, color='blue', edgecolor='black', stat='probability' # 此参数控制区间概率和为1 ) # 绘制参考曲线 x = np.linspace(0.01, 20, 1000) y = 0.5 * np.exp(-0.5 * x) plt.plot(x, y, lw=3, c='r', label='Chi sqrd with df=2') plt.legend(loc='upper right') plt.show()
如果用matplotlib原生函数,可手动将权重归一化到总和为1后绘制:
# 归一化权重,使总和为1 norm_weights = weights / weights.sum() plt.hist( vals, bins=num_bins, weights=norm_weights, color='blue', edgecolor='black' ) # 参考曲线绘制同上 plt.legend(loc='upper right') plt.show()
3. 若需概率密度(积分和为1)
如果要对比概率密度曲线,只需把stat参数改为'density'(seaborn),或在matplotlib中设置density=True即可:
sns.histplot( vals, weights=weights, bins=num_bins, color='blue', edgecolor='black', stat='density' ) plt.plot(x, y, lw=3, c='r', label='Chi sqrd with df=2') plt.legend(loc='upper right') plt.show()
内容的提问来源于stack exchange,提问作者q2w3e4
相关产品推荐
相关产品推荐

