使用matplotlib同X轴双Y轴绘制两个pandas时间序列出现异常
异常原因
核心原因是两个DataFrame的时间索引完全无重叠,触发了pandas绘图的索引对齐逻辑:
- df1的索引为不含周末的工作日,df2的所有索引都是周日,二者没有公共的日期值
- pandas的
DataFrame.plot默认优先按索引标签对齐数据,再映射到x轴位置:- 第一个子图先绘制df1,再绘制df2时,df2的所有日期都不在df1的索引范围内,对应值全部被填充为NaN,因此看不到第二条曲线
- 第二个子图先绘制df2,再绘制df1时,df1的31个日期无任何能匹配df2的索引,pandas会强行将df1的所有数据挤压到df2仅有的5个x轴刻度位置上,就出现了异常的重采样错位现象
解决方案
提供两种可直接复用的方案:
方案1:使用matplotlib原生绘图接口(推荐)
直接指定x、y参数绘图,不会触发pandas的索引对齐逻辑,完全按实际时间戳位置渲染:
import matplotlib.pyplot as plt import numpy as np import pandas as pd # 原数据构造代码保持不变 dates1 = ['2021-08-26', '2021-08-27', '2021-08-30', '2021-08-31', '2021-09-01', '2021-09-02', '2021-09-03', '2021-09-07', '2021-09-08', '2021-09-09', '2021-09-10', '2021-09-13', '2021-09-14', '2021-09-15', '2021-09-16', '2021-09-17', '2021-09-20', '2021-09-21', '2021-09-22', '2021-09-23', '2021-09-24', '2021-09-27', '2021-09-28', '2021-09-29', '2021-09-30', '2021-10-01', '2021-10-04', '2021-10-05', '2021-10-06', '2021-10-07', '2021-10-08'] dates2 = ['2021-08-29', '2021-09-05', '2021-09-12', '2021-09-19', '2021-09-26'] y1 = np.random.randn(len(dates1)).cumsum() y2 = np.random.randn(len(dates2)).cumsum() df1 = pd.DataFrame({'date':pd.to_datetime(dates1), 'y1':y1}) df1.set_index('date', inplace=True) df2 = pd.DataFrame({'date':pd.to_datetime(dates2), 'y2':y2}) df2.set_index('date', inplace=True) # 绘图代码修改为如下内容 fig, axs = plt.subplots(1,4, figsize=[12,4]) # 子图0 双Y轴 axs[0].plot(df1.index, df1['y1'], label='y1') ax_twin0 = axs[0].twinx() ax_twin0.plot(df2.index, df2['y2'], color='orange', label='y2') # 合并图例 lines1, labels1 = axs[0].get_legend_handles_labels() lines2, labels2 = ax_twin0.get_legend_handles_labels() axs[0].legend(lines1 + lines2, labels1 + labels2, loc='upper left') # 子图1 双Y轴 axs[1].plot(df2.index, df2['y2'], label='y2') ax_twin1 = axs[1].twinx() ax_twin1.plot(df1.index, df1['y1'], color='orange', label='y1') # 合并图例 lines1, labels1 = axs[1].get_legend_handles_labels() lines2, labels2 = ax_twin1.get_legend_handles_labels() axs[1].legend(lines1 + lines2, labels1 + labels2, loc='upper left') # 子图2、3保持原有逻辑 df1.y1.plot(ax=axs[2]) df2.y2.plot(ax=axs[3]) plt.tight_layout() plt.show()
方案2:先合并数据集再绘图
将两个DataFrame按日期外连接合并为一个完整数据集,保证索引覆盖所有日期后再用pandas绘图:
# 数据构造完成后先执行合并 df_all = pd.concat([df1, df2], axis=1).sort_index() # 绘图代码修改为如下内容 fig, axs = plt.subplots(1,4, figsize=[12,4]) df_all['y1'].plot(ax=axs[0]) df_all['y2'].plot(ax=axs[0], secondary_y=True) df_all['y2'].plot(ax=axs[1]) df_all['y1'].plot(ax=axs[1], secondary_y=True) df1.y1.plot(ax=axs[2]) df2.y2.plot(ax=axs[3]) plt.tight_layout() plt.show()
内容的提问来源于stack exchange,提问作者Alberto
相关产品推荐
相关产品推荐

