You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何并行化groupby后的筛选与排序操作?

并行化GroupBy后的筛选与排序操作

核心思路

把循环内的单组处理逻辑封装成独立函数,借助Python的多进程库并行处理所有分组,最后统一合并结果,替代原串行循环的低效方式。

步骤与代码实现

  1. 封装单组处理函数
    抽离原循环内的筛选、排序、取数逻辑,处理单个分组并返回有效结果(空分组返回None):

    def process_group(group_data):
        _, group = group_data
        # 合并筛选条件,减少数据遍历次数
        group = group[(group['intersection_area'] >= 80) & (group['clear_percent'] >= 10)]
        if group.empty:
            return None
        # 按clear_percent降序排序后取第一条(修正原代码未赋值、未指定排序方向的问题)
        group = group.sort_values(['clear_percent'], ascending=False).head(1)
        return group
    
  2. 并行处理所有分组
    使用concurrent.futures.ProcessPoolExecutor实现并行计算,最后用pd.concat合并结果(比循环append效率提升显著):

    import pandas as pd
    import time
    from concurrent.futures import ProcessPoolExecutor
    
    # 假设df为已加载的原始数据
    groups = df.groupby(['date'])
    group_list = list(groups)  # 将分组转换为可迭代的列表
    
    tt = time.time()
    # 启动进程池并行处理
    with ProcessPoolExecutor() as executor:
        results = executor.map(process_group, group_list)
    
    # 过滤空结果并合并为最终DataFrame
    new_df = pd.concat([res for res in results if res is not None], ignore_index=True)
    ids = new_df['scene_id'].to_list()
    
    print(f"并行处理耗时: {time.time() - tt:.2f}秒")
    

关键注意点

  • 排序逻辑修正:原代码中sort_values未赋值且未指定排序方向,若要取clear_percent最高的记录,必须添加ascending=False并重新赋值给group。
  • 并行收益前提:仅当分组数量多、单组数据量大时,并行处理才有性能优势;小数据量下进程启动的开销可能抵消提速效果。
  • Windows系统适配:使用多进程库时,需把主逻辑放在if __name__ == '__main__':代码块内,避免子进程重复初始化的报错:
    if __name__ == '__main__':
        # 此处放置df定义、groupby、并行处理等主代码
        pass
    

内容的提问来源于stack exchange,提问作者Manap Shymyr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 06:24:31