You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何异步文件夹遍历比同步顺序遍历速度更慢?

问题:异步文件夹遍历比同步版本慢的原因?

我正在用aiofiles和标准os库做文件夹遍历的基准测试,以下是测试代码:

异步版本函数

import os
import time
from typing import List
import aiofiles.os as aio_os
import asyncio

subfolders1= set()
async def traverrse_local_path() -> List[str]:  

    async def traverse(local_path, count=None):
        
        try:
            current_dirs_list = [local_path + '/' + f for f in os.listdir(local_path)]
            current_subfolders = [f for f in current_dirs_list if os.path.isdir(f)]

            if not current_subfolders:
                if any([await aio_os.path.isfile(f) for f in current_dirs_list]):
                    subfolders1.add(local_path)
                    return
            for p in current_subfolders:
                    await traverse(p, count)
        except WindowsError:
            pass

    start = time.perf_counter()
    path_to_proj = 'C:/Program Files'
    list_folders = [path_to_proj + '/' + f for f in  os.listdir(path_to_proj) if os.path.isdir(path_to_proj + '/' + f)]
    await asyncio.gather(*(traverse(f,count) for count,f in enumerate(list_folders)))
    end = time.perf_counter()
    print(end-start)

loop = asyncio.get_event_loop()
try:
    loop.run_until_complete(traverrse_local_path())
finally:
    loop.close()
print(len(subfolders1))

同步顺序版本函数

subfolders_2 = set()
def traverrse_local_path() -> List[str]:
    
    def traverse(local_path, count=None):
        try:
            current_dirs_list = [local_path + '/' + f for f in  os.listdir(local_path)]
            current_subfolders = [f for f in current_dirs_list if os.path.isdir(f)]
            if not current_subfolders:
                if any([os.path.isfile(f) for f in current_dirs_list]):
                    subfolders_2.add(local_path)
                    return
            for p in current_subfolders:
                    traverse(p, count)
        except WindowsError:
            pass

    start = time.perf_counter()
    path_to_proj = 'C:/Program Files'
    traverse(path_to_proj)
    end = time.perf_counter()
    print(end - start)

traverrse_local_path()
print(len(subfolders_2))

时间对比结果

时间对比图

我原本以为异步版本应该更快,但实际结果却相反,请问这是为什么?


原因分析
  • 异步操作的额外开销:每一次await都会触发协程切换,这个过程需要保存和恢复协程状态,本身存在开销。同步代码直接调用系统函数,没有这些调度成本,在磁盘IO等待时间不足以抵消切换开销的场景下,异步版本自然更慢。

  • 未真正利用异步IO优势:你的异步代码里,os.listdir、os.path.isdir都是同步调用,仅最后判断文件时用了异步的aio_os.path.isfile。核心遍历操作还是同步逻辑,异步只是套了一层协程壳,反而增加了复杂度和额外开销。要发挥异步优势,得把所有文件系统操作替换为aiofiles的异步版本,比如await aio_os.listdir(local_path)、await aio_os.path.isdir(f)。

  • 磁盘IO的并发瓶颈:不管是机械硬盘还是SSD,磁盘物理结构决定了它无法同时高效处理大量并行IO请求。异步并发发起多个遍历操作,会导致磁盘频繁寻址,反而降低整体效率;而同步顺序遍历能让磁盘更高效地连续读取数据。

  • 协程调度的累积开销:你用asyncio.gather同时启动多个根目录的遍历协程,这些协程执行中会频繁切换,累积的调度成本进一步拉低了异步版本的性能。

内容的提问来源于stack exchange,提问作者PilotChalkanov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 16:33:37