Pandas DataFrame唯一值总数统计柱状图:求高效内置方法
替换低效自定义类别求和方法的Pandas内置方案
我需要按类别统计数值总和(以颜色为类别示例),虽然已经实现了功能,但自定义的combine_like_categories方法效率低下、扩展性差,希望用Python库的内置函数替代。以下是我的原始实现代码:
import pandas as pd import matplotlib.pyplot as plt import numpy as np # Sum all values for each unique category. def combine_like_categories(df): col = df['color'] new_colors = col.unique() new_values = np.zeros_like(new_colors) new_df = pd.DataFrame(np.array([new_colors, new_values]).T) headers=['color', 'value'] new_df.columns = headers for _, row in df.iterrows(): new_df_index = new_df.loc[new_df['color']==row['color']].index[0] new_df.iloc[new_df_index, 1] += row['value'] return new_df data = [ ['red', 3], ['blue', 2], ['green', 5], ['orange', 3], ['blue', 1], ['red', 7] ] headers=['color', 'value'] df = pd.DataFrame(data) df.columns = headers hist_df = combine_like_categories(df) plt.figure() plt.bar(hist_df['color'], hist_df['value']) plt.title('Counts of Each Color') plt.xlabel('Color') plt.ylabel('Counts') plt.show()
运行代码后生成对应柱状图:
解决方案:使用Pandas内置的groupby+sum
你完全可以用Pandas原生的分组聚合方法替代自定义函数,这能解决效率和扩展性问题:
修改后的核心代码
只需要一行代码就能替代整个combine_like_categories函数:
hist_df = df.groupby('color')['value'].sum().reset_index()
完整优化后代码
import pandas as pd import matplotlib.pyplot as plt data = [ ['red', 3], ['blue', 2], ['green', 5], ['orange', 3], ['blue', 1], ['red', 7] ] headers=['color', 'value'] df = pd.DataFrame(data, columns=headers) # 用Pandas内置分组聚合替代自定义函数 hist_df = df.groupby('color')['value'].sum().reset_index() plt.figure() plt.bar(hist_df['color'], hist_df['value']) plt.title('各颜色数值总和') plt.xlabel('颜色') plt.ylabel('数值总和') plt.show()
优势说明
- 效率大幅提升:
groupby是Pandas优化过的向量化操作,避免了iterrows的逐行循环,数据量越大,性能差距越明显 - 扩展性极强:如果需要更换分类列(比如换成其他类别字段)、调整聚合逻辑(比如求均值
mean()、计数count()),只需修改groupby参数和聚合方法即可 - 代码更简洁可读:一行代码完成原本自定义函数的全部逻辑,维护成本更低
内容的提问来源于stack exchange,提问作者greenerpastures
相关产品推荐
相关产品推荐

