如何用Pandas在单张图中对比多年逐小时流量数据?
没问题!针对你的流量数据对比绘图需求,我整理了几个实用的方案,用Pandas搭配Matplotlib/Seaborn就能轻松实现,一起来看看:
先做数据预处理(可选但推荐)
首先确保你的时间索引是标准的datetime格式,同时提取年份、月份等信息,方便后续分组对比:
import pandas as pd import matplotlib.pyplot as plt import seaborn as sns # 确保索引为datetime类型(如果原始数据还不是的话) year.index = pd.to_datetime(year.index) # 新增年份、月份、以及「月-日-时」列(用于同一月份内的时间对齐) year['year'] = year.index.year year['month'] = year.index.month year['day_hour'] = year.index.strftime('%m-%d %H:00')
方案1:对比特定年份的同一月份小时级流量
比如你想直观对比2016年1月和2017年1月的每小时流量变化,可以用折线图展示:
# 筛选出1月的数据 jan_data = year[year['month'] == 1] # 把年份转成列,方便绘图时对比 jan_pivot = jan_data.pivot(index='day_hour', columns='year', values='discharge') # 绘制对比折线图 plt.figure(figsize=(12, 6)) jan_pivot.plot(ax=plt.gca()) plt.title('Hourly Discharge Comparison: January 2016 vs January 2017') plt.xlabel('Date & Hour') plt.ylabel('Discharge (m^3/s)') plt.xticks(rotation=45) # 旋转x轴标签避免拥挤 plt.legend(title='Year') plt.tight_layout() # 自动调整布局 plt.show()
方案2:对比所有年份的同一月份日均流量
如果小时级数据太密集,你可以先聚合为日均流量,对比每年同月的整体趋势:
# 筛选1月数据,按年份+日期分组计算日均流量 daily_jan = year[year['month'] == 1].groupby(['year', year.index.date])['discharge'].mean().unstack(level=0) # 绘制日均流量对比图 plt.figure(figsize=(12, 6)) daily_jan.plot(ax=plt.gca()) plt.title('Daily Average Discharge Comparison: January Across Years') plt.xlabel('Date in January') plt.ylabel('Average Discharge (m^3/s)') plt.xticks(rotation=45) plt.legend(title='Year') plt.tight_layout() plt.show()
方案3:箱线图对比不同月份的流量分布
如果想对比多个月份的流量整体分布(比如1月和7月的流量差异),箱线图是不错的选择:
# 筛选要对比的月份,比如1月和7月 compare_months = year[year['month'].isin([1, 7])] # 绘制箱线图 plt.figure(figsize=(8, 6)) sns.boxplot(data=compare_months, x='month', y='discharge', hue='year') plt.title('Discharge Distribution Comparison: January vs July') plt.xlabel('Month') plt.ylabel('Discharge (m^3/s)') plt.tight_layout() plt.show()
小提示:
- 要对比其他月份,只需要修改代码中的月份数字(比如
month == 2就是2月,isin([3,9])就是对比3月和9月) - 如果数据量极大,小时级折线图会显得拥挤,可以考虑进一步聚合为3小时平均或者日均值
- 你可以通过
plt.style.use('seaborn-v0_8')或者调整Seaborn的参数来自定义图表风格
内容的提问来源于stack exchange,提问作者Johan R
相关产品推荐
相关产品推荐

