如何在Pandas中把不同日期的时间序列绘制在同一图表并共用坐标轴?
解决方案
核心思路是剥离日期信息,只保留一天内的时间维度,让两个数据集共享同一个24小时的X轴,具体步骤如下:
1. 数据预处理:提取一天内的时间
首先将两个数据集的时间列转换为datetime类型,然后提取时分秒(或转换为一天内的总秒数)作为统一的X轴依据。
假设你的数据是Pandas DataFrame格式,示例代码:
import pandas as pd # 读取数据(示例) df_may = pd.read_csv('may2_data.csv') df_jun = pd.read_csv('jun3_data.csv') # 转换为datetime类型 df_may['timestamp'] = pd.to_datetime(df_may['timestamp']) df_jun['timestamp'] = pd.to_datetime(df_jun['timestamp']) # 方式1:提取时分秒(离散时间点) df_may['time_of_day'] = df_may['timestamp'].dt.time df_jun['time_of_day'] = df_jun['timestamp'].dt.time # 方式2:转换为一天内的总秒数(连续数值轴,更适合绘图) df_may['total_seconds'] = df_may['timestamp'].dt.hour * 3600 + df_may['timestamp'].dt.minute * 60 + df_may['timestamp'].dt.second df_jun['total_seconds'] = df_jun['timestamp'].dt.hour * 3600 + df_jun['timestamp'].dt.minute * 60 + df_jun['timestamp'].dt.second
2. (可选)对齐时间点(解决时间点差异问题)
如果两个数据集的时间点差异较大,直接绘图会出现曲线断点或错位,可通过重采样或插值对齐时间粒度:
# 重采样到1分钟粒度(用均值填充缺失值,可根据需求替换为sum/median等) df_may_resampled = df_may.set_index('timestamp').resample('1min').mean().reset_index() df_jun_resampled = df_jun.set_index('timestamp').resample('1min').mean().reset_index() # 重新提取连续秒数作为X轴 df_may_resampled['total_seconds'] = df_may_resampled['timestamp'].dt.hour * 3600 + df_may_resampled['timestamp'].dt.minute * 60 + df_may_resampled['timestamp'].dt.second df_jun_resampled['total_seconds'] = df_jun_resampled['timestamp'].dt.hour * 3600 + df_jun_resampled['timestamp'].dt.minute * 60 + df_jun_resampled['timestamp'].dt.second
3. 绘制共享XY轴的对比图
使用Matplotlib绘制两条曲线,共享同一组X/Y轴:
import matplotlib.pyplot as plt plt.figure(figsize=(12, 6)) # 用连续秒数作为X轴(推荐,曲线更平滑) plt.plot(df_may['total_seconds'], df_may['value'], label='5月2日', alpha=0.7) plt.plot(df_jun['total_seconds'], df_jun['value'], label='6月3日', alpha=0.7) # 设置X轴刻度为小时,提升可读性 plt.xticks(ticks=[i*3600 for i in range(0, 25)], labels=[f'{i}:00' for i in range(0, 25)]) plt.xlabel('一天中的时间') plt.ylabel('数值') plt.title('两天24小时时间序列对比') plt.legend() # 调整布局避免标签被截断 plt.tight_layout() plt.show()
注意事项
- 如果使用时分秒作为X轴,由于是离散的时间对象,绘图时可能出现刻度密集的情况,可通过
plt.xticks(rotation=45)旋转标签解决。 - 若数据存在缺失值,可在重采样后用
interpolate()方法进行线性插值,让曲线更连贯:df_may_resampled['value'] = df_may_resampled['value'].interpolate()
内容的提问来源于stack exchange,提问作者redcode
相关产品推荐
相关产品推荐

