Matplotlib子图中相同数据直方图存在差异的问题排查
直方图子图显示不一致问题排查与解决
我尝试为稀疏数据集绘制直方图,因单图显示近乎空白,故将数据划分为四个不同区间制作子图以提升可读性。但发现相同数据区间的两个子图直方图存在差异,尽管二者基于完全相同的数据与参数。无法分享原始数据集,但已生成模拟数据复现该问题,代码如下:
import numpy as np import scipy.stats as st import matplotlib import matplotlib.pyplot as plt dist = st.chi.rvs(0.233966, loc=-5.29021*10**-23, scale=3.955368*10**4, size=110000) bins = np.concatenate((np.linspace(0, 10000, 1000,), np.linspace(10100,100000, 1000), np.linspace(101000, 1000000, 1000), np.linspace(1010000, 4000000, 40))) fig, axs = plt.subplots(ncols=2, nrows=2, figsize=(12,8)) y, x, l = axs[0,0].hist(dist, bins=bins, density=True) start_plot = 0 end_plot = 5000 y_max = np.max(y[(x[:-1]<end_plot) & (x[:-1]>start_plot)]) axs[0, 0].set_xlim((start_plot, end_plot)) axs[0,0].set_ylim((0, y_max)) axs[0,1].hist(dist, bins=bins, density=True) start_plot = 0 end_plot = 5000 y_max = np.max(y[(x[:-1]<end_plot) & (x[:-1]>start_plot)]) axs[0,1].set_xlim((start_plot, end_plot)) axs[0,1].set_ylim((0, y_max)) axs[1,0].hist(dist, bins=bins, density=True) start_plot = 5000 end_plot = 50000 y_max = np.max(y[(x[:-1]<end_plot) & (x[:-1]>start_plot)]) axs[1,0].set_xlim((start_plot, end_plot)) axs[1,0].set_ylim((0, y_max)) axs[1,1].hist(dist, bins=bins, density=True) start_plot = 50000 end_plot = 1000000 y_max = np.max(y[(x[:-1]<end_plot) & (x[:-1]>start_plot)]) axs[1,1].set_xlim((start_plot, end_plot)) axs[1,1].set_ylim((0, y_max)) plt.show()
顶部两个子图应显示相同数据,但差异明显;若不绘制底部两个子图,顶部子图虽更相似但仍存在不同。查看补丁发现其x位置、宽度和高度均一致,现求助该差异产生的原因及解决方法。
问题原因
这是Matplotlib自带的自动坐标轴对齐特性导致的视觉错觉:
- 当绘制多个子图时,Matplotlib会自动调整每个子图的刻度与轴范围,尽可能让不同子图的刻度线对齐,保证整体布局的整齐性。但这会让相同数据的子图在视觉上出现差异,实际直方图的每个柱子(补丁)的位置、宽度、高度都是完全一致的。
- 即使只绘制顶部两个子图,Matplotlib仍会进行细微的刻度调整,因此视觉上还是会存在微小差别。
解决方法
方法1:关闭自动坐标轴对齐
在创建子图时关闭constrained_layout,手动调整子图间距:
fig, axs = plt.subplots(ncols=2, nrows=2, figsize=(12,8), constrained_layout=False) plt.subplots_adjust(wspace=0.3, hspace=0.3) # 手动设置子图间的宽、高间距
方法2:手动同步子图坐标轴属性
对于需要完全一致显示的子图,直接复制第一个子图的坐标轴范围、刻度等属性:
# 绘制完axs[0,0]后,将其坐标轴属性同步给axs[0,1] axs[0,1].set_xlim(axs[0,0].get_xlim()) axs[0,1].set_ylim(axs[0,0].get_ylim()) axs[0,1].set_xticks(axs[0,0].get_xticks()) axs[0,1].set_yticks(axs[0,0].get_yticks())
方法3:预先计算直方图数据,统一绘制
先一次性计算好直方图的数值和区间,再用bar方法分别绘制到每个子图,避免重复调用hist可能带来的细微差异(虽然本例中计算结果一致,但此方法更稳妥):
# 预先计算直方图数据 y, x = np.histogram(dist, bins=bins, density=True) widths = np.diff(x) # 绘制第一个子图 axs[0,0].bar(x[:-1], y, width=widths) start_plot = 0 end_plot = 5000 y_max = np.max(y[(x[:-1]<end_plot) & (x[:-1]>start_plot)]) axs[0,0].set_xlim((start_plot, end_plot)) axs[0,0].set_ylim((0, y_max)) # 绘制第二个子图,复用同一组数据 axs[0,1].bar(x[:-1], y, width=widths) axs[0,1].set_xlim(axs[0,0].get_xlim()) axs[0,1].set_ylim(axs[0,0].get_ylim()) axs[0,1].set_xticks(axs[0,0].get_xticks()) axs[0,1].set_yticks(axs[0,0].get_yticks()) # 后续子图同理...
内容的提问来源于stack exchange,提问作者Av06
相关产品推荐
相关产品推荐

