You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python Multiprocessor并行化For循环

Python HTTP请求循环并行化改造方案

原有代码并行失效的核心原因

你之前写的ThreadPoolExecutor调用没有生效,是因为executor.submit(post_query(x))这行直接在主线程执行了post_query(x),把函数执行后的返回值(None)传给了submit,线程池根本没有接收到可调度的任务,所有请求还是串行执行的。

另外你的原函数把循环逻辑写在了函数内部,无法拆分单个任务给并行单元调度,首先要把单个请求的逻辑抽离为独立函数。

推荐方案:多线程实现(适配IO密集型场景)

发HTTP请求属于典型的IO密集型任务,大部分耗时在等待网络响应,用多线程的调度开销远小于多进程,是这个场景的最优选择。

  1. 先抽离单任务执行函数:
import json
import requests
from concurrent.futures import ThreadPoolExecutor, as_completed

def post_single_query(question_id):
    url = f"{global_url}card/{str(question_id)}/query"
    payload = json.dumps({"ignore_cache": False})
    headers = {
        "Content-Type": "application/json",
        "Cookie": f"metabase.SESSION={session}"
    }
    resp = requests.post(url, headers=headers, data=payload)
    return resp.text
  1. 正确的线程池调用写法:
if __name__ == "__main__":
    questions_ids = get_cards()  # 替换为你的ID列表来源
    # max_workers根据接口承载能力调整,IO场景一般设为5-20即可,不要开过高避免触发限流
    with ThreadPoolExecutor(max_workers=8) as executor:
        # 正确传参:传入函数引用+参数,不要加括号直接调用函数
        futures = [executor.submit(post_single_query, qid) for qid in questions_ids]
        # 按任务完成顺序获取返回结果
        for future in as_completed(futures):
            try:
                print(future.result())
            except Exception as e:
                print(f"请求执行出错: {str(e)}")

如果需要返回结果和传入的ID顺序保持一致,可以用更简洁的map写法:

if __name__ == "__main__":
    questions_ids = get_cards()
    with ThreadPoolExecutor(max_workers=8) as executor:
        for resp_text in executor.map(post_single_query, questions_ids):
            print(resp_text)

可选方案:多进程实现(不推荐本场景使用)

多进程更适合CPU密集型计算任务,对于HTTP请求这类IO等待场景,多进程启动开销大、资源占用高,效率不如多线程。如果确实需要用multiprocessing类方案,写法和线程池几乎一致,注意必须放在if __name__ == "__main__":块下执行,避免跨平台运行异常:

from concurrent.futures import ProcessPoolExecutor

if __name__ == "__main__":
    questions_ids = get_cards()
    # 多进程并发数不要超过CPU核心数
    with ProcessPoolExecutor(max_workers=4) as executor:
        futures = [executor.submit(post_single_query, qid) for qid in questions_ids]
        for future in as_completed(futures):
            print(future.result())

注意事项

  • 不要把并发数设置过高,否则容易触发目标接口的限流策略,甚至压垮目标服务
  • 建议加异常捕获逻辑,避免单个请求失败导致整个并行任务中断
  • 如果全局变量global_url、session在子进程/子线程中存在读取异常,直接作为参数传入post_single_query函数即可

内容的提问来源于stack exchange,提问作者Felipe Abello

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 02:06:31