You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何并行化生成pandas DataFrame的for循环并合并结果以缩短处理耗时

Pandas年度DataFrame生成循环并行化方案

核心思路:每个make_df(year)的调用互相独立无依赖,可先并行生成所有年份的DataFrame,再一次性合并到初始表中,大幅降低耗时占比最高的函数计算环节总用时。

实现代码

import pandas as pd
from concurrent.futures import ProcessPoolExecutor
from functools import reduce

# 初始表保持不变
df = pd.DataFrame({'key': ['K0', 'K1', 'K2', 'K3'],
                      'A': ['A0', 'A1', 'A2', 'A3']})

# 你的复杂逻辑函数无需修改,直接保留
def make_df(year):
    df = pd.DataFrame({'key': ['K0', 'K1', 'K2', 'K3'], str(year): [str(year), str(year+1), str(year+2), str(year+3)]})
    return df

if __name__ == '__main__':
    # 定义需要处理的年份列表
    years = list(range(2020, 2015, -1))
    # 并行执行所有make_df任务
    with ProcessPoolExecutor() as executor:
        year_dfs = list(executor.map(make_df, years))
    # 一次性合并所有生成的年度表
    df = reduce(lambda left, right: pd.merge(left, right, on='key', how='left'), year_dfs, df)

说明

  • 耗时优化逻辑:原来串行循环总用时为所有make_df调用耗时之和,并行后总用时基本等于单次最慢的make_df调用耗时,函数逻辑越复杂、年份越多,提升效果越明显。
  • CPU密集型任务优先用ProcessPoolExecutor,可规避Python GIL锁限制;IO密集型任务(如读文件、调接口)优先用ThreadPoolExecutor,进程开销更低。
  • Windows系统必须把并行执行逻辑放到if __name__ == '__main__'块内,否则会出现进程启动报错。
  • 合并环节为纯内存操作,速度极快,不会成为性能瓶颈,无需额外并行。

内容的提问来源于stack exchange,提问作者JHP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 03:24:01