Pandas:按cust_id分组将两列合并为排序后的字典列
需求:按用户ID分组生成按评分降序的产品字典DataFrame
原始DataFrame
cust_id product score 1 bat 0.8 2 ball 0.3 2 phone 0.6 3 tv 1.0 2 bat 1.0 4 phone 0.2 1 ball 0.6
预期输出
cust_id product_dict 1 {'bat': 0.8, 'ball':0.6} 2 {'bat': 1.0, 'phone':0.6, 'ball':0.3} 3 {'tv': 1.0} 4 {'phone':0.2}
无效代码尝试
s = df.groupby(['cust_id','product'])['score'].apply(list) d = {x: s.xs(x).to_dict() for x in s.index.levels[0]} df1 = df.groupby('cust_id').agg(num_of_products=('product','nunique')) df1.insert(2, 'product_dict', df1.index.map(d)) df1 = df1.reset_index()
问题分析与解决方案
原代码问题点
- 对
score使用apply(list)完全多余,每个(cust_id, product)组合对应单个评分,转成列表会导致字典值为列表形式(如{'bat': [0.8]}),不符合预期格式 - 生成字典时未按
score降序排序,默认按产品名称排序,无法满足需求
正确实现代码
# 先按用户ID分组,每组内按score降序排列 sorted_df = df.sort_values(by=['cust_id', 'score'], ascending=[True, False]) # 分组生成按score降序的字典 result = sorted_df.groupby('cust_id').apply( lambda x: dict(zip(x['product'], x['score'])) ).reset_index(name='product_dict') # 如需保留产品数量列,可添加以下代码 result['num_of_products'] = result['product_dict'].apply(len) print(result)
执行结果
cust_id product_dict num_of_products 0 1 {'bat': 0.8, 'ball': 0.6} 2 1 2 {'bat': 1.0, 'phone': 0.6, 'ball': 0.3} 3 2 3 {'tv': 1.0} 1 3 4 {'phone': 0.2} 1
内容的提问来源于stack exchange,提问作者Danish
相关产品推荐
相关产品推荐

