Pandas aggregate()生成异常DataFrame,绘图报错求助
问题
使用Pandas处理CSV文件,通过groupby()和aggregate()聚合数据后,生成的DataFrame中id列显示形式特殊,无法直接用于matplotlib绘图。
教程代码:
import pandas as pd df = pd.DataFrame({'id': [101, 101, 102, 103, 103, 103], 'employee': ['Dan', 'Dan', 'Rick', 'Ken', 'Ken', 'Ken'], 'sales': [4, 1, 3, 2, 5, 3], 'returns': [1, 2, 2, 1, 3, 2]}) agg_functions = {'employee': 'first', 'sales': 'sum', 'returns': 'sum'} df_new = df.groupby(df['id']).aggregate(agg_functions)
生成的DataFrame:
employee sales returns id 101 Dan 5 3 102 Rick 3 2 103 Ken 10 6
尝试绘图时报错:
# 报错代码1 ax = df_new.plot(kind="line", x=df_new["id"], y=df_new["sales"])
报错信息:
KeyError: 'id'
# 报错代码2 ax = df_new.plot(kind="line", x=df_new.index, y=df_new["sales"])
报错信息:
KeyError: "None of [Index([101, 102, 103], dtype='int64')] are in the [columns]"
原因
groupby()默认会把分组字段设为DataFrame的索引,而非普通列。所以df_new中id是索引,不是列,直接用df_new["id"]会找不到对应列;而df.plot()的x参数默认接受列名,传入索引对象也会报错。
解决方法
方法1:聚合时保留id为普通列
在groupby()中添加as_index=False参数,让分组字段保持为列:
df_new = df.groupby('id', as_index=False).aggregate(agg_functions)
此时id是普通列,直接用列名绘图即可:
ax = df_new.plot(kind="line", x='id', y='sales')
方法2:将现有索引转为列
如果已经生成了带索引的df_new,用reset_index()把索引转为普通列:
df_new = df_new.reset_index()
之后同样可以用x='id'绘图。
方法3:直接使用索引绘图(无需转列)
df.plot()默认会用索引作为x轴,或者直接指定索引名称作为x参数(这里索引名为id):
# 方式1:默认用索引做x轴,只指定y列 ax = df_new.plot(kind="line", y='sales') # 方式2:指定x为索引名称 ax = df_new.plot(kind="line", x='id', y='sales')
提取
id值的方法 - 若
id是索引:# 获取索引对象 id_index = df_new.index # 获取值数组 id_values = df_new.index.values - 若
id已转为列:id_series = df_new['id'] id_values = df_new['id'].values
内容的提问来源于stack exchange,提问作者Rerun
相关产品推荐
相关产品推荐

