Matplotlib绘制跨年度数据时2023年数据被压缩的问题排查
问题描述
我有一个用于绘制6个月时间序列图表的Python函数,代码如下:
def result(data): fig = plt.figure(figsize=(12, 8), dpi=150) ax = plt.subplot(111) data = data.sort_values(by='period', ascending=1) box = ax.get_position() ax.spines['top'].set_visible(True) ax.spines['right'].set_visible(True) plt.plot(data['period'], data['result'], label='Result', solid_capstyle='round', marker='o', color='#92c6ff') plt.plot(data['period'], data['peer_result'], label='Peer Result', color='#5990C6', linestyle='--') annotate_chart(data['period'], data['result'].round(1)) plt.legend(loc='upper center', fontsize=20, bbox_to_anchor=(0.5, -0.15), frameon=False, ncol=4) plt.ylabel('(%)', fontsize=28, weight='bold') plt.xticks(data['period'], data['month'], rotation=45, fontsize=28) plt.yticks(fontsize=28) plt.title('Results', fontsize=28) plt.tight_layout() img = files_loc + '/result.png' plt.savefig(img) plt.close(fig) # Line fill graph placement pdf.set_xy(4, 30) pdf.image(img, w=105, h=75)
对应的数据集格式:
| period | month | result | peer_result |
|---|---|---|---|
| 202308 | AUG | 82.14 | 83.22 |
| 202309 | SEP | 81.47 | 82.37 |
| 202310 | OCT | 82.87 | 82.53 |
| 202311 | NOV | 84.68 | 83.52 |
| 202312 | DEC | 83.26 | 82.58 |
| 202401 | JAN | 83.28 | 82.78 |
遇到的问题:
- 传入完整数据集(含2024年1月)时,图表中2023年的数据被压缩,x轴间距异常
- 仅传入2023年5条数据时,图表显示正常
- 传入
data.tail(5)(含202310到202401)时问题再次出现,说明和数据集大小无关 - 尝试将
period转为数值型或分类型均无效
问题根源
核心问题是x轴用了period的数值/字符串值作为坐标,但跨年后的数值间距和同年内完全不一致:
- 同年内月份的
period数值差是1(比如202308到202309差1) - 202312到202401的数值差是89(202401-202312=89)
Matplotlib会严格按照数值间距分配x轴空间,导致前面的月份挤成一团,最后一个月份占了大量空间。就算转成分类类型,只要没把它当成等间隔的时间序列处理,问题依然存在。
解决方法
把period转换成真正的datetime类型,让Matplotlib识别这是等间隔的时间轴:
1. 转换period为datetime类型
在函数开头或数据预处理阶段添加转换代码:
data['period'] = pd.to_datetime(data['period'], format='%Y%m')
2. 调整标注函数(如果需要)
如果annotate_chart函数依赖period的数值定位标注,得改成用数据的位置索引(比如range(len(data))),因为datetime类型的x轴坐标是时间戳,直接用会导致标注位置偏移。
修改后的完整函数
def result(data): fig = plt.figure(figsize=(12, 8), dpi=150) ax = plt.subplot(111) # 转换为datetime类型,让Matplotlib识别时间序列 data['period'] = pd.to_datetime(data['period'], format='%Y%m') data = data.sort_values(by='period', ascending=1) box = ax.get_position() ax.spines['top'].set_visible(True) ax.spines['right'].set_visible(True) plt.plot(data['period'], data['result'], label='Result', solid_capstyle='round', marker='o', color='#92c6ff') plt.plot(data['period'], data['peer_result'], label='Peer Result', color='#5990C6', linestyle='--') # 用位置索引替代原period数值,适配datetime轴的标注 annotate_chart(range(len(data)), data['result'].round(1)) plt.legend(loc='upper center', fontsize=20, bbox_to_anchor=(0.5, -0.15), frameon=False, ncol=4) plt.ylabel('(%)', fontsize=28, weight='bold') plt.xticks(data['period'], data['month'], rotation=45, fontsize=28) plt.yticks(fontsize=28) plt.title('Results', fontsize=28) plt.tight_layout() img = files_loc + '/result.png' plt.savefig(img) plt.close(fig) # Line fill graph placement pdf.set_xy(4, 30) pdf.image(img, w=105, h=75)
这样处理后,所有月份的x轴间距会被自动设置为等间隔,彻底解决数据压缩问题。
内容的提问来源于stack exchange,提问作者Agata
相关产品推荐
相关产品推荐

