基于多列的Pandas多聚合函数实现(非分层索引)
Pandas聚合:按Color和Shape计算Count与Sum(无分层索引)
刚好我之前也处理过类似的需求,用pivot_table完全可以实现,而且很容易避免分层索引,给你一步步拆解:
首先先构造你给出的输入DataFrame:
import pandas as pd # 构造输入数据 data = { 'Color': ['Blue', 'Red', 'Green', 'Blue', 'Blue', 'Green', 'Red', 'Blue', 'Blue'], 'Shape': ['Square', 'Square', 'Square', 'Circle', 'Square', 'Circle', 'Circle', 'Square', 'Circle'], 'Value': [5, 2, 7, 9, 2, 6, 2, 5, 1] } df = pd.DataFrame(data)
接下来用pivot_table实现聚合,同时处理掉分层索引:
# 使用pivot_table完成聚合,避免分层索引 agg_result = pd.pivot_table( df, index=['Color', 'Shape'], # 按Color和Shape分组 values='Value', # 聚合的目标列 aggfunc={'Value': ['count', 'sum']} # 同时计算计数和求和 ).reset_index() # 把行索引中的Color、Shape转为普通列 # 重命名列名,匹配你需要的格式 agg_result.columns = ['Color', 'Shape', 'Count', 'Sum'] # 查看最终结果 print(agg_result)
运行这段代码后,输出结果就和你期望的完全一致:
Color Shape Count Sum 0 Blue Circle 2 10 1 Blue Square 3 12 2 Green Circle 1 6 3 Green Square 1 7 4 Red Circle 1 2 5 Red Square 1 2
关键说明:
pivot_table默认会生成分层列索引(比如('Value', 'count')这种),我们通过reset_index()把行索引里的分组字段变回普通列,再手动重命名列名就能消除分层结构。- 如果想要结果的顺序和你给出的完全一致,可以在最后加上
sort_values(by=['Color', 'Shape'], ascending=[True, False])来调整排序。
内容的提问来源于stack exchange,提问作者pacificdune
相关产品推荐
相关产品推荐

