You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于直方图创建密度图?验证方法并实现平滑曲线

密度图实现疑问与平滑密度图绘制方法

现有直方图代码及结果

我现有如下绘制直方图的脚本:

import matplotlib.pyplot as plt
import numpy as np

fig5 = plt.figure()
fig5.subplots_adjust(wspace=0.4, hspace=0.6)
ax5_2 = fig5.add_subplot(212)
std_error = "{:.2f}".format(np.std(df['error']))
ave_error = "{:.2f}".format(np.nanmean(df['error']))
ax5_2.set_title("Error Histogram ave:" + ave_error + " std:" + std_error)
ax5_2.hist(df['error'], range=(0, 2), bins=20)

运行后得到误差值在0-2区间的柱状直方图,标题包含误差的均值与标准差。

尝试的密度图代码及问题

我希望在此基础上绘制密度图,尝试了以下代码:

hist,bin_edges=np.histogram(df['error'], range=(0, 2), bins=20,density=True)
bin_widths=np.diff(bin_edges)
densities = hist / (np.sum(hist) * bin_widths)
plt.bar(bin_edges[:-1], densities, width=bin_widths, align='edge')

# Customize your plot as needed
plt.xlabel('X-axis Label')
plt.ylabel('Density')
plt.title('Density Plot')

得到的是柱状形式的密度图,但我认为密度图应该是平滑折线图,同时不确定这个实现是否正确。

另外我曾尝试用Seaborn绘制,但得到的直方图和密度图似乎完全不相关。

补充测试的差异情况

我尝试了如下Seaborn代码:

errors = np.random.randn(100)
df = pd.DataFrame({'errors': errors})
sns.histplot(data=df, x="errors", kde=True)

得到的图中直方图与密度曲线匹配。但当我用Matplotlib做等价操作:

fig5 = plt.figure()
fig5.subplots_adjust(wspace=0.4, hspace=0.6)
ax5_2 = fig5.add_subplot(212)
ax5_2.hist(errors, range=(-3, 3), bins=9)

得到手动设置区间和分箱的直方图,再运行自己写的密度图代码:

hist,bin_edges=np.histogram(errors, range=(-3, 3), bins=9,density=True)
bin_widths=np.diff(bin_edges)
densities = hist / (np.sum(hist) * bin_widths)
plt.bar(bin_edges[:-1], densities, width=bin_widths, align='edge')

两者结果差异很大。


问题解答

1. 你的密度图实现是否正确?

不正确。当np.histogram设置density=True时,返回的hist已经是归一化的密度值(满足曲线下面积为1),不需要再执行densities = hist / (np.sum(hist) * bin_widths)这一步——该操作会错误地二次缩放数值,导致结果严重偏离预期。

2. 如何绘制对应直方图的平滑折线密度图?

提供两种可靠方法:

方法一:Seaborn对齐参数绘制(最简单)

Seaborn的histplot可以直接叠加平滑核密度曲线,只需将参数与你的Matplotlib直方图对齐(指定range和bins),就能保证两者匹配:

import seaborn as sns
import pandas as pd
import numpy as np

# 替换为你的真实数据df['error']
errors = np.random.randn(100)
df = pd.DataFrame({'error': errors})

# 对齐直方图参数:区间、分箱数
sns.histplot(data=df, x="error", kde=True, bins=9, range=(-3, 3))

方法二:Matplotlib+Scipy手动绘制(无依赖Seaborn)

用Scipy的gaussian_kde计算核密度估计,再在原有直方图上叠加平滑折线:

import matplotlib.pyplot as plt
import numpy as np
from scipy.stats import gaussian_kde

# 你的原始直方图代码
fig5 = plt.figure()
fig5.subplots_adjust(wspace=0.4, hspace=0.6)
ax5_2 = fig5.add_subplot(212)
std_error = "{:.2f}".format(np.std(df['error']))
ave_error = "{:.2f}".format(np.nanmean(df['error']))
ax5_2.set_title(f"Error Histogram ave:{ave_error} std:{std_error}")
# 绘制直方图并获取密度值(density=True)
n, bins, patches = ax5_2.hist(df['error'], range=(0,2), bins=20, density=True, alpha=0.6)

# 计算平滑密度曲线
valid_errors = df['error'].dropna()  # 剔除NaN值
kde = gaussian_kde(valid_errors)
xvals = np.linspace(0, 2, 1000)  # 与直方图区间保持一致
density_curve = kde(xvals)

# 绘制平滑折线
ax5_2.plot(xvals, density_curve, 'r-', label='Density')
ax5_2.legend()
plt.show()

3. 测试结果差异的原因

你之前用Seaborn时未指定range和bins,Seaborn会自动选择区间和分箱数,与你手动用Matplotlib设置的range=(-3,3), bins=9不一致,因此看起来“不相关”。只要将Seaborn的参数与Matplotlib对齐,就能得到匹配的直方图和密度曲线。


内容的提问来源于stack exchange,提问作者KansaiRobot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 20:19:51