如何用Pandas按combo_id聚合行并生成字典键值对?
Pandas按combo_id聚合生成字典的解决方法
原始数据
| id1 | id2 | attr1 | combo_id | perm_id |
|---|---|---|---|---|
| 1 | 2 | [9606] | [1,2] | AB |
| 2 | 1 | [9606] | [1,2] | BA |
| 3 | 4 | [9606] | [3,4] | AB |
| 4 | 3 | [9606] | [3,4] | BA |
需求
将combo_id相同的行聚合,以每行的perm_id为键、attr1为值生成字典,最终得到如下结果:
| attr1 | combo_id |
|---|---|
| {'AB':[9606], 'BA': [9606]} | [1,2] |
| {'AB':[9606], 'BA': [9606]} | [3,4] |
你遇到的问题
你尝试先将attr1转为字典:
df['attr1'] = df.apply(lambda x: {x['perm_id']: x['attr1']})
然后合并同组字典:
df.groupby(['combo_id']).agg({'attr1': lambda x: {x**}})
但出现KeyError: perm_id错误。
错误原因&解决步骤
1. 错误根源
df.apply()默认按列处理(axis=0),此时lambda里的x是整列Series,不是单行数据,所以找不到perm_id这个键。必须指定axis=1让函数逐行处理。
2. 正确实现方式
方法一:先转换列再聚合
先修正apply的axis参数,再用聚合函数合并字典:
# 逐行生成单个键值对的字典 df['attr1'] = df.apply(lambda x: {x['perm_id']: x['attr1']}, axis=1) # 分组后合并所有字典 result = df.groupby('combo_id')['attr1'].agg(lambda x: {k: v for d in x for k, v in d.items()}).reset_index()
方法二:直接在聚合中完成(更高效)
不需要先修改原DataFrame,直接在groupby的agg里处理每行数据并合并:
result = df.groupby('combo_id').agg( attr1=lambda x: {row['perm_id']: row['attr1'] for _, row in x.iterrows()} ).reset_index()
两种方法都能得到你需要的结果,方法二避免修改原数据,更推荐。
内容的提问来源于stack exchange,提问作者Jia Wu
相关产品推荐
相关产品推荐

