如何用Pandas按ID分组并将多列合并为列表/字典
按ID分组将多列合并为嵌套列表的实现方法
原始数据表
| ID | Component | Revenue |
|---|---|---|
| 1 | 4 | 10 |
| 1 | 5 | 20 |
| 2 | 4 | 15 |
| 3 | 6 | 30 |
预期输出
| ID | Grouped |
|---|---|
| 1 | [[4, 10], [5, 20]] |
| 2 | [4, 15] |
| 3 | [6, 30] |
已知单列分组转列表可通过以下代码实现:
df.groupby(['ID']).Component.apply(list).reset_index()
要实现多列合并为嵌套列表的需求,可以用以下两种方法:
方法一:使用groupby+apply逐行打包
import pandas as pd # 构造示例数据 df = pd.DataFrame({ 'ID': [1, 1, 2, 3], 'Component': [4, 5, 4, 6], 'Revenue': [10, 20, 15, 30] }) # 分组生成嵌套列表 result = df.groupby('ID').apply( lambda x: [list(row) for row in x[['Component', 'Revenue']].values] ).reset_index(name='Grouped') # 处理单元素组的格式,将外层列表展开 result['Grouped'] = result['Grouped'].apply(lambda x: x[0] if len(x) == 1 else x) print(result)
方法二:使用agg结合zip打包列
import pandas as pd df = pd.DataFrame({ 'ID': [1, 1, 2, 3], 'Component': [4, 5, 4, 6], 'Revenue': [10, 20, 15, 30] }) # 分组打包列数据 result = df.groupby('ID').agg( Grouped=lambda x: list(zip(x['Component'], x['Revenue'])) ).reset_index() # 将元组转为列表,并处理单元素组格式 result['Grouped'] = result['Grouped'].apply( lambda x: [list(item) for item in x] if len(x) > 1 else list(x[0]) ) print(result)
两种方法都能输出符合预期的结果,其中处理单元素组的步骤是为了让单个条目从[[4,15]]转为[4,15],如果不需要这个格式转换,可以去掉对应apply语句。
内容的提问来源于stack exchange,提问作者John Miller
相关产品推荐
相关产品推荐

