You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python线程写入文件为何不比同步IO更快?

Python多线程写文件为何没提速?

你的测试结果很正常,原因既不是GIL导致同步执行,也不是操作系统只能处理单个文件写入,核心问题出在磁盘IO的物理特性和多线程的额外开销上,具体拆解:

1. GIL不背这个锅

Python的GIL确实会限制CPU密集型任务的多线程并行,但文件写入是操作系统级的IO操作——当Python调用file.open()或json.dump()时,会主动释放GIL,让其他线程可以执行。所以你的多线程写入过程中,线程是真的在并行发起IO请求的。

2. 磁盘是真正的瓶颈

不管是机械硬盘(HDD)还是固态硬盘(SSD),都有天生的IO限制:

  • HDD:磁头同一时间只能在一个磁道上读写,多线程写不同文件时,磁头需要频繁切换位置(寻道),反而会增加额外开销,总耗时反而可能比单线程顺序写入更长。
  • SSD:虽然支持并行读写,但单块SSD的总带宽是有限的。如果你的单线程写入已经接近磁盘带宽上限,多线程也没法提速;如果写入量小,多线程的调度开销(线程切换、任务分配)反而会抵消并行带来的收益,导致总耗时持平甚至变慢。

3. 你的代码场景验证

你的代码是给每个response_data单独写一个id.json文件,属于多文件写入场景:

  • 若用HDD,磁头频繁在不同文件的存储位置间切换,效率极低;
  • 若用SSD,单线程写入的速度已经足够利用磁盘带宽,多线程带来的额外开销会让总时间没优势。

一些优化思路

  • 如果是HDD:尽量合并写入或者按磁盘存储顺序批量写,减少磁头寻道次数;
  • 如果是SSD:可以尝试调整线程池大小(比如设置max_workers等于CPU核心数),但提升幅度可能有限;
  • 避免在系统盘(根目录/)进行大量写入,系统盘通常有其他进程占用IO资源,会进一步影响性能。

你的测试代码:

def write_file(response_data: dict) -> None:
    try:
        file_path = Path('/', f"{response_data['id']}.json")
        with file_path.open("w") as file:
            json.dump(response_data, file)
    except Exception as e:
        log.error("Unable to write file")
        log.error("".join(traceback.TracebackException.from_exception(e).format()))

log.info("Time start threads")
start = time.time()
with concurrent.futures.ThreadPoolExecutor() as executor:
    executor.map(write_file, final_responses)
log.info("Time end threads")
end = time.time()
log.info("Total time for threads %s", (end - start) / 60)

log.info("Time start no threads")
start = time.time()
for response in final_responses:
    write_file(response)
log.info("Time end no threads")
end = time.time()
log.info("Total time for no threads %s", (end - start) / 60)

内容的提问来源于stack exchange,提问作者Matthew Barlowe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 17:24:15