Pandas按多列分组后计算字典元素中位数并生成新列的实现方法
实现代码
import pandas as pd import statistics # 构造示例DataFrame,你可以替换为自己的数据源 data = [ ["ItemA", 0, "p", {"store1":50,"store2":70,"store3":90,"store4":44,"store5":76}], ["ItemB", 0, "p", {"store2":22,"store3":15,"store4":77,"store5":0}], ["ItemC", 0, "p", {"store1":46,"store2":13,"store3":9,"store4":87,"store5":45}], ["ItemD", 0, "q", {"store1":88,"store2":16,"store4":5,"store5":2}], ["ItemE", 0, "q", {"store1":7,"store2":55}], ["ItemF", 1, "t", {"store3":25,"store4":75,"store5":87}], ["ItemG", 1, "t", {"store1":32,"store3":66,"store4":87,"store5":0}], ["ItemH", 1, "t", {"store1":54,"store2":33,"store3":12,"store4":67,"store5":8}], ] df = pd.DataFrame(data, columns=["item", "category", "subcategory", "sales_count"]) # 核心计算逻辑 def calc_group_median(group_sales): # 展平当前分组内所有sales_count字典的数值 all_sales_values = [val for sales_dict in group_sales for val in sales_dict.values()] return statistics.median(all_sales_values) # 按双字段分组计算中位数,直接映射回原表行 df["median_across_group"] = df.groupby(["category", "subcategory"])["sales_count"].transform(calc_group_median)
结果验证
各个分组最终计算得到的中位数如下:
category=0, subcategory=p:所有销售值排序后为[0,9,13,15,22,44,45,46,50,70,76,77,87,90],中位数为45.5category=0, subcategory=q:所有销售值排序后为[2,5,7,16,55,88],中位数为11.5category=1, subcategory=t:所有销售值排序后为[0,8,12,25,32,33,54,66,67,75,87,87],中位数为43.5
内容的提问来源于stack exchange,提问作者charlie_boy
相关产品推荐
相关产品推荐

