Python pandas调用groupby()分组后绘图无法按指定字段分色怎么办
DataFrame分组后按分类字段分色绘图方案
现有数据样例
待处理的DataFrame结构如下:
Processed Card Transaction ID Transaction amount Error_Occured Date 2019-01-01 Carte Rouge 217142203412 147924.21 0 2019-01-01 ChinaPay 149207925233 65301.63 1 2019-01-01 Masterkard 766507067450 487356.91 5 2019-01-01 VIZA 145484139636 97774.52 1 2019-01-02 Carte Rouge 510466748547 320951.10 3
目标效果
X轴为Date字段,Y轴为Error_Occured字段,数据点/折线按Processed Card字段区分颜色
原有代码问题
之前尝试的代码如下,运行后会绘制Transaction ID等无关字段的系列,无法按Processed Card分色:
df = df.groupby(['Date','Processed Card']).sum('Error_Occured') df = df.reset_index() df.set_index("Date",inplace=True) df.plot(legend=True) plt.show()
问题出在两点:
- 分组聚合时未指定仅聚合
Error_Occured列,导致Transaction ID、Transaction amount这类无关数值列也被纳入聚合结果,被pandas默认识别为待绘制的Y轴系列 - 未将长表数据转换为适合pandas原生plot的宽表结构,也未指定分类映射字段,无法按Processed Card的取值拆分系列
可直接运行的实现代码
方法1:pandas原生实现(无需额外依赖)
import pandas as pd import matplotlib.pyplot as plt # 分组聚合仅保留需要的Error_Occured字段 df_agg = df.groupby(['Date', 'Processed Card'], as_index=False)['Error_Occured'].sum() # 透视为宽表:行索引为日期,列对应不同支付卡,值为错误数 df_pivot = df_agg.pivot(index='Date', columns='Processed Card', values='Error_Occured') # 绘图,加marker参数可同时显示数据点 df_pivot.plot(legend=True, marker='o') plt.ylabel('Error_Occured') plt.show()
方法2:seaborn实现(代码更简洁,无需手动透视)
import seaborn as sns import matplotlib.pyplot as plt # 仅聚合需要的字段 df_agg = df.groupby(['Date', 'Processed Card'], as_index=False)['Error_Occured'].sum() # 直接指定x、y轴和分色字段hue即可 sns.lineplot(data=df_agg, x='Date', y='Error_Occured', hue='Processed Card', marker='o') plt.show()
内容的提问来源于stack exchange,提问作者MaximeTars
相关产品推荐
相关产品推荐

