如何高效分组DataFrame并同时获取分组键与对应列值用于函数调用?
问题描述
我有一个包含"time"和"location"列的DataFrame,示例如下:
data = [['A', '13:00', 2], ['B', '15:00', 0], ['C', '13:00', 3], ['D', '13:00', 0], ['E', '17:00', 0]] df = pd.DataFrame(data, columns=['location', 'time', 'other'])
我希望按time对数据分组,之后调用一个需要传入time和对应locations列表的函数。我尝试了以下代码:
result = df.groupby('time')['location'].apply(list)
但此时time会作为result的索引被隐藏,我无法以向量化方式同时将time和locations传入函数。result的结构如下:
13:00 [A, C, D] 15:00 [B] 17:00 [E]
当我迭代该结果或使用apply()时,只能获取到locations列表,而我需要将time也作为函数参数。请问正确的实现方法是什么?
解决方案
方法1:重置索引将分组键转为列
先把分组后的索引(time)转为普通列,这样就能同时获取时间和对应的地点列表:
# 分组后重置索引,指定列表列的名称 grouped = df.groupby('time')['location'].apply(list).reset_index(name='locations') # 定义需要调用的示例函数 def process_data(time_str, locations): return f"时段{time_str}涵盖地点:{', '.join(locations)}" # 逐行应用函数,同时传入time和locations grouped['processed_result'] = grouped.apply( lambda row: process_data(row['time'], row['locations']), axis=1 )
处理后的grouped结构如下:
| time | locations | processed_result |
|---|---|---|
| 13:00 | ['A', 'C', 'D'] | 时段13:00涵盖地点:A, C, D |
| 15:00 | ['B'] | 时段15:00涵盖地点:B |
| 17:00 | ['E'] | 时段17:00涵盖地点:E |
方法2:直接在分组apply中获取分组键
对groupby对象直接应用函数时,可以通过group.name拿到当前分组的time值,同时获取分组对应的子DataFrame:
def process_group(group_key, group_df): locations = group_df['location'].tolist() return f"时段{group_key}涵盖地点:{', '.join(locations)}" # 对分组对象应用函数,自动传入分组键和子DataFrame result = df.groupby('time').apply(lambda g: process_group(g.name, g))
执行后result的输出为:
time 13:00 时段13:00涵盖地点:A, C, D 15:00 时段15:00涵盖地点:B 17:00 时段17:00涵盖地点:E dtype: object
方法3:聚合时保留分组键为列
使用agg聚合时设置as_index=False,让time保持为普通列,一步到位得到包含time和locations的DataFrame:
# 聚合时指定as_index=False,避免time成为索引 grouped = df.groupby('time', as_index=False).agg(locations=('location', list)) # 应用函数的方式和方法1一致 grouped['processed_result'] = grouped.apply( lambda row: process_data(row['time'], row['locations']), axis=1 )
内容的提问来源于stack exchange,提问作者meelszz
相关产品推荐
相关产品推荐

