pandas groupby求均值后sort_values报by参数意外错误如何解决
问题原因
- 你分组聚合后得到的
gdm是Series结构,而非DataFrame:你是对单列Stars做分组求均值,最终结果的索引是Country,值为对应国家的Stars均值,本身没有名为Country的列。 - Series的
sort_values()方法没有by参数,该参数仅用于DataFrame指定排序参考列,这是触发TypeError的直接原因。 - 即使不触发参数报错,你指定
by='Country'也不符合需求:你的目标是按计算得到的均值降序排序,而非按国家名称排序。 - 额外错误:
ascending参数需要传入布尔值False,你写的字符串'False'无法实现降序效果。
修正代码
如果你习惯用Series操作,直接修改排序和绘图逻辑即可:
import pandas as pd import matplotlib.pyplot as plt f2 = pd.read_csv('ramen-ratings.csv') # 先转换Stars为数值类型 f2['Stars'] = pd.to_numeric(f2['Stars'], errors='coerce') # 按Country分组求均值,得到索引为Country、值为平均Stars的Series gdm = f2.groupby('Country')['Stars'].mean() # 排序:Series直接调用sort_values,不需要by参数,传布尔值False实现降序 gd_sorted = gdm.sort_values(ascending=False) # 绘图 plt.figure(figsize = (12,8)) plt.bar(gd_sorted.index, gd_sorted.values) # 可选:x轴标签旋转避免文字重叠 plt.xticks(rotation=90) plt.xlabel('国家') plt.ylabel('平均评分') plt.show()
如果你习惯用DataFrame操作,可以在聚合后转成DataFrame再排序:
gdm = f2.groupby('Country')['Stars'].mean().reset_index() # 此时为DataFrame结构,可以用by参数指定按平均Stars列降序 gd_sorted = gdm.sort_values(by='Stars', ascending=False) # 绘图时可以直接传入列名 plt.bar('Country', 'Stars', data=gd_sorted)
内容的提问来源于stack exchange,提问作者Jibby Ola
相关产品推荐
相关产品推荐

