You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas:按cust_id分组将两列合并为排序后的字典列

需求:按用户ID分组生成按评分降序的产品字典DataFrame

原始DataFrame

cust_id   product       score
1         bat           0.8
2         ball          0.3
2         phone         0.6
3         tv            1.0
2         bat           1.0
4         phone         0.2
1         ball          0.6 

预期输出

cust_id   product_dict       
1         {'bat': 0.8, 'ball':0.6}         
2         {'bat': 1.0, 'phone':0.6, 'ball':0.3}         
3         {'tv': 1.0}
4         {'phone':0.2}

无效代码尝试

s = df.groupby(['cust_id','product'])['score'].apply(list)

d = {x: s.xs(x).to_dict() for x in s.index.levels[0]}

df1 = df.groupby('cust_id').agg(num_of_products=('product','nunique'))

df1.insert(2, 'product_dict', df1.index.map(d))
df1 = df1.reset_index()

问题分析与解决方案

原代码问题点

  1. 对score使用apply(list)完全多余,每个(cust_id, product)组合对应单个评分,转成列表会导致字典值为列表形式(如{'bat': [0.8]}),不符合预期格式
  2. 生成字典时未按score降序排序,默认按产品名称排序,无法满足需求

正确实现代码

# 先按用户ID分组,每组内按score降序排列
sorted_df = df.sort_values(by=['cust_id', 'score'], ascending=[True, False])

# 分组生成按score降序的字典
result = sorted_df.groupby('cust_id').apply(
    lambda x: dict(zip(x['product'], x['score']))
).reset_index(name='product_dict')

# 如需保留产品数量列,可添加以下代码
result['num_of_products'] = result['product_dict'].apply(len)

print(result)

执行结果

cust_id                          product_dict  num_of_products
0        1                {'bat': 0.8, 'ball': 0.6}                2
1        2  {'bat': 1.0, 'phone': 0.6, 'ball': 0.3}                3
2        3                          {'tv': 1.0}                1
3        4                        {'phone': 0.2}                1

内容的提问来源于stack exchange,提问作者Danish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 15:16:01