Python 3.9嵌套循环并行化求助:如何用多进程缩短代码耗时
优化Python代码并行执行方案(针对IO密集型网络请求)
你的代码核心瓶颈是多层嵌套的串行网络请求,结合你已经用了orjson/ujson和requests会话的优化,下面给出基于Python 3.9的多进程(及更适合的多线程)实现方案:
核心思路
所有getJson调用属于IO密集型任务,这类任务的等待时间远大于CPU计算时间,因此:
- 若坚持用多进程,使用
concurrent.futures.ProcessPoolExecutor - 更推荐用多线程(
ThreadPoolExecutor),因为线程切换开销远低于进程,且IO等待时GIL会自动释放,效率更高
以下是具体实现代码:
1. 基础工具函数优化
import requests import orjson from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor # 优化后的JSON获取函数,自带错误处理 def get_json(url): # 每个任务独立创建Session(多进程/线程无法共享Session) with requests.Session() as session: try: response = session.get(url, timeout=10) response.raise_for_status() # 触发HTTP错误(如404/500) return orjson.loads(response.text) except requests.exceptions.RequestException as e: print(f"请求失败 {url}: {str(e)}") return None
2. 并行化重构waitforEvent函数
方案一:多进程实现
# 单独封装获取单场比赛详情的函数 def fetch_single_match(match_id): return get_json(f"{match_api}{match_id}") # 单独封装获取单个赛事系列的所有比赛详情 def fetch_series_matches(series_id): series_url = f"{series_api}{series_id}" series_json = get_json(series_url) if not series_json: return [] match_ids = [match["id"] for match in series_json["series"]["matches"]] # 内部用线程池并行获取该系列下的比赛(IO密集型用线程更高效) with ThreadPoolExecutor(max_workers=10) as thread_exec: return list(thread_exec.map(fetch_single_match, match_ids)) def waitforEvent(event_id): # 先获取基础赛事数据(这里假设json_page是对应event_id的URL,若不是请自行修正) event_json = get_json(json_page) if not event_json: return [] series_ids = [s["id"] for s in event_json["event"]["series"]] # 多进程并行处理所有赛事系列 with ProcessPoolExecutor() as proc_exec: # ProcessPoolExecutor默认max_workers为CPU核心数,可根据需求调整 all_match_details = list(proc_exec.map(fetch_series_matches, series_ids)) # 扁平化结果(可选,根据后续处理需求调整) return [match for series_matches in all_match_details for match in series_matches]
方案二:全线程实现(更推荐)
如果不需要处理CPU密集型任务,全线程方案开销更低、速度更快:
def waitforEvent(event_id): event_json = get_json(json_page) if not event_json: return [] # 收集所有需要请求的比赛URL all_match_urls = [] for s in event_json["event"]["series"]: series_json = get_json(f"{series_api}{s['id']}") if series_json: all_match_urls.extend([f"{match_api}{match['id']}" for match in series_json["series"]["matches"]]) # 批量并行请求所有比赛详情 with ThreadPoolExecutor(max_workers=20) as exec: return list(exec.map(get_json, all_match_urls))
关键注意事项
- Session隔离:多进程/线程间不能共享requests.Session,必须每个任务独立创建,避免连接池混乱
- 并发数控制:根据目标API的限流规则调整
max_workers,过大的并发数可能导致被封禁 - 错误处理:务必添加异常捕获,避免单个请求失败导致整个并行任务崩溃
- 变量修正:原代码中
json_page未关联event_id,请确保该URL是对应event_id的正确地址
内容的提问来源于stack exchange,提问作者Alazin
相关产品推荐
相关产品推荐

