如何基于直方图创建密度图?验证方法并实现平滑曲线
现有直方图代码及结果
我现有如下绘制直方图的脚本:
import matplotlib.pyplot as plt import numpy as np fig5 = plt.figure() fig5.subplots_adjust(wspace=0.4, hspace=0.6) ax5_2 = fig5.add_subplot(212) std_error = "{:.2f}".format(np.std(df['error'])) ave_error = "{:.2f}".format(np.nanmean(df['error'])) ax5_2.set_title("Error Histogram ave:" + ave_error + " std:" + std_error) ax5_2.hist(df['error'], range=(0, 2), bins=20)
运行后得到误差值在0-2区间的柱状直方图,标题包含误差的均值与标准差。
尝试的密度图代码及问题
我希望在此基础上绘制密度图,尝试了以下代码:
hist,bin_edges=np.histogram(df['error'], range=(0, 2), bins=20,density=True) bin_widths=np.diff(bin_edges) densities = hist / (np.sum(hist) * bin_widths) plt.bar(bin_edges[:-1], densities, width=bin_widths, align='edge') # Customize your plot as needed plt.xlabel('X-axis Label') plt.ylabel('Density') plt.title('Density Plot')
得到的是柱状形式的密度图,但我认为密度图应该是平滑折线图,同时不确定这个实现是否正确。
另外我曾尝试用Seaborn绘制,但得到的直方图和密度图似乎完全不相关。
补充测试的差异情况
我尝试了如下Seaborn代码:
errors = np.random.randn(100) df = pd.DataFrame({'errors': errors}) sns.histplot(data=df, x="errors", kde=True)
得到的图中直方图与密度曲线匹配。但当我用Matplotlib做等价操作:
fig5 = plt.figure() fig5.subplots_adjust(wspace=0.4, hspace=0.6) ax5_2 = fig5.add_subplot(212) ax5_2.hist(errors, range=(-3, 3), bins=9)
得到手动设置区间和分箱的直方图,再运行自己写的密度图代码:
hist,bin_edges=np.histogram(errors, range=(-3, 3), bins=9,density=True) bin_widths=np.diff(bin_edges) densities = hist / (np.sum(hist) * bin_widths) plt.bar(bin_edges[:-1], densities, width=bin_widths, align='edge')
两者结果差异很大。
问题解答
1. 你的密度图实现是否正确?
不正确。当np.histogram设置density=True时,返回的hist已经是归一化的密度值(满足曲线下面积为1),不需要再执行densities = hist / (np.sum(hist) * bin_widths)这一步——该操作会错误地二次缩放数值,导致结果严重偏离预期。
2. 如何绘制对应直方图的平滑折线密度图?
提供两种可靠方法:
方法一:Seaborn对齐参数绘制(最简单)
Seaborn的histplot可以直接叠加平滑核密度曲线,只需将参数与你的Matplotlib直方图对齐(指定range和bins),就能保证两者匹配:
import seaborn as sns import pandas as pd import numpy as np # 替换为你的真实数据df['error'] errors = np.random.randn(100) df = pd.DataFrame({'error': errors}) # 对齐直方图参数:区间、分箱数 sns.histplot(data=df, x="error", kde=True, bins=9, range=(-3, 3))
方法二:Matplotlib+Scipy手动绘制(无依赖Seaborn)
用Scipy的gaussian_kde计算核密度估计,再在原有直方图上叠加平滑折线:
import matplotlib.pyplot as plt import numpy as np from scipy.stats import gaussian_kde # 你的原始直方图代码 fig5 = plt.figure() fig5.subplots_adjust(wspace=0.4, hspace=0.6) ax5_2 = fig5.add_subplot(212) std_error = "{:.2f}".format(np.std(df['error'])) ave_error = "{:.2f}".format(np.nanmean(df['error'])) ax5_2.set_title(f"Error Histogram ave:{ave_error} std:{std_error}") # 绘制直方图并获取密度值(density=True) n, bins, patches = ax5_2.hist(df['error'], range=(0,2), bins=20, density=True, alpha=0.6) # 计算平滑密度曲线 valid_errors = df['error'].dropna() # 剔除NaN值 kde = gaussian_kde(valid_errors) xvals = np.linspace(0, 2, 1000) # 与直方图区间保持一致 density_curve = kde(xvals) # 绘制平滑折线 ax5_2.plot(xvals, density_curve, 'r-', label='Density') ax5_2.legend() plt.show()
3. 测试结果差异的原因
你之前用Seaborn时未指定range和bins,Seaborn会自动选择区间和分箱数,与你手动用Matplotlib设置的range=(-3,3), bins=9不一致,因此看起来“不相关”。只要将Seaborn的参数与Matplotlib对齐,就能得到匹配的直方图和密度曲线。
内容的提问来源于stack exchange,提问作者KansaiRobot

