使用DataFrame聚合生成字典:Pandas学习遇到的分组问题
Hey there! I've got you covered on this Pandas aggregation task. Let's walk through the solution step by step.
First, let's make sure we're working with the same sample DataFrame:
import pandas as pd # Your original data df = pd.DataFrame({ 'id': [1, 2, 1, 2], 'batchid': [101, 101, 102, 104], 'batchname': ['accounts', 'accounts', 'finance', 'admin'] })
The issue with your apply() + stack() approach is that stack() reshapes the DataFrame into a long format, which doesn't help us build the list of dictionaries we need. Instead, we can use groupby() paired with to_dict('records')—this method converts each row in a group into a dictionary, and we just need to collect those dictionaries into a list per id.
Here's the working code:
# Group by 'id', then convert each group's batch rows to a list of dictionaries result_df = df.groupby('id').apply( lambda group: group[['batchid', 'batchname']].to_dict('records') ).reset_index(name='group_info')
Let's check the output:
id group_info 0 1 [{'batchid': 101, 'batchname': 'accounts'}, {'batchid': 102, 'batchname': 'finance'}] 1 2 [{'batchid': 101, 'batchname': 'accounts'}, {'batchid': 104, 'batchname': 'admin'}]
Perfect, that's exactly the output you wanted! The to_dict('records') method is key here—it creates a list of dictionaries where each dictionary represents a row in the grouped subset, with column names as keys and row values as values.
If you want to avoid lambda for better readability, you can also define a small helper function:
def batch_to_dict_list(group): return group[['batchid', 'batchname']].to_dict('records') result_df = df.groupby('id').apply(batch_to_dict_list).reset_index(name='group_info')
This works exactly the same way, just a bit more explicit.
内容的提问来源于stack exchange,提问作者deepa warrier

