You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3.9嵌套循环并行化求助:如何用多进程缩短代码耗时

优化Python代码并行执行方案(针对IO密集型网络请求)

你的代码核心瓶颈是多层嵌套的串行网络请求,结合你已经用了orjson/ujson和requests会话的优化,下面给出基于Python 3.9的多进程(及更适合的多线程)实现方案:

核心思路

所有getJson调用属于IO密集型任务,这类任务的等待时间远大于CPU计算时间,因此:

  • 若坚持用多进程,使用concurrent.futures.ProcessPoolExecutor
  • 更推荐用多线程(ThreadPoolExecutor),因为线程切换开销远低于进程,且IO等待时GIL会自动释放,效率更高

以下是具体实现代码:

1. 基础工具函数优化

import requests
import orjson
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor

# 优化后的JSON获取函数,自带错误处理
def get_json(url):
    # 每个任务独立创建Session(多进程/线程无法共享Session)
    with requests.Session() as session:
        try:
            response = session.get(url, timeout=10)
            response.raise_for_status()  # 触发HTTP错误(如404/500)
            return orjson.loads(response.text)
        except requests.exceptions.RequestException as e:
            print(f"请求失败 {url}: {str(e)}")
            return None

2. 并行化重构waitforEvent函数

方案一:多进程实现

# 单独封装获取单场比赛详情的函数
def fetch_single_match(match_id):
    return get_json(f"{match_api}{match_id}")

# 单独封装获取单个赛事系列的所有比赛详情
def fetch_series_matches(series_id):
    series_url = f"{series_api}{series_id}"
    series_json = get_json(series_url)
    if not series_json:
        return []
    
    match_ids = [match["id"] for match in series_json["series"]["matches"]]
    # 内部用线程池并行获取该系列下的比赛(IO密集型用线程更高效)
    with ThreadPoolExecutor(max_workers=10) as thread_exec:
        return list(thread_exec.map(fetch_single_match, match_ids))

def waitforEvent(event_id):
    # 先获取基础赛事数据(这里假设json_page是对应event_id的URL,若不是请自行修正)
    event_json = get_json(json_page)
    if not event_json:
        return []
    
    series_ids = [s["id"] for s in event_json["event"]["series"]]
    
    # 多进程并行处理所有赛事系列
    with ProcessPoolExecutor() as proc_exec:
        # ProcessPoolExecutor默认max_workers为CPU核心数,可根据需求调整
        all_match_details = list(proc_exec.map(fetch_series_matches, series_ids))
    
    # 扁平化结果(可选,根据后续处理需求调整)
    return [match for series_matches in all_match_details for match in series_matches]

方案二:全线程实现(更推荐)

如果不需要处理CPU密集型任务,全线程方案开销更低、速度更快:

def waitforEvent(event_id):
    event_json = get_json(json_page)
    if not event_json:
        return []
    
    # 收集所有需要请求的比赛URL
    all_match_urls = []
    for s in event_json["event"]["series"]:
        series_json = get_json(f"{series_api}{s['id']}")
        if series_json:
            all_match_urls.extend([f"{match_api}{match['id']}" for match in series_json["series"]["matches"]])
    
    # 批量并行请求所有比赛详情
    with ThreadPoolExecutor(max_workers=20) as exec:
        return list(exec.map(get_json, all_match_urls))

关键注意事项

  • Session隔离:多进程/线程间不能共享requests.Session,必须每个任务独立创建,避免连接池混乱
  • 并发数控制:根据目标API的限流规则调整max_workers,过大的并发数可能导致被封禁
  • 错误处理:务必添加异常捕获,避免单个请求失败导致整个并行任务崩溃
  • 变量修正:原代码中json_page未关联event_id,请确保该URL是对应event_id的正确地址

内容的提问来源于stack exchange,提问作者Alazin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 07:37:47