如何使用Pandas的Groupby函数按年份分组统计不同燃油类型的占比
占比计算实现代码
你之前使用transform未成功大概率是因为没有指定以year为单位分组计算年度总车辆数,以下是完整可运行的实现:
# 1. 统计每年各燃油类型的车辆数量 year_fuel = my_data_frame.groupby(['year', 'fuelType']).size().reset_index(name='counts') # 2. 计算每一年的总车辆数,transform会将分组求和结果广播到对应年份的每一行 year_fuel['year_total'] = year_fuel.groupby('year')['counts'].transform('sum') # 3. 计算占比,保留2位小数 year_fuel['percentage'] = (year_fuel['counts'] / year_fuel['year_total'] * 100).round(2) # 4. 可选:生成和你示例格式一致的分层索引结果 result = year_fuel.set_index(['year', 'fuelType'])[['percentage']] print(result)
输出结果和你要求的示例格式完全匹配。
堆叠柱状图绘制代码
要绘制横轴为年份的堆叠占比图,只需要将数据转为宽表格式后直接调用pandas的绘图接口即可:
import matplotlib.pyplot as plt # 转为绘图所需的宽表,缺失值(某年份无对应燃油类型数据)填充为0 plot_data = year_fuel.pivot(index='year', columns='fuelType', values='percentage').fillna(0) # 绘制堆叠柱状图 plot_data.plot( kind='bar', stacked=True, figsize=(14,7), title='奥迪二手车燃油类型占比年度变化趋势' ) plt.xlabel('年份') plt.ylabel('占比(%)') plt.legend(title='燃油类型') plt.xticks(rotation=45) plt.tight_layout() plt.show()
内容的提问来源于stack exchange,提问作者JacobMarlo
相关产品推荐
相关产品推荐

