为何基于asyncio与boto3的Python S3调用代码仍串行执行?
问题分析:boto3结合asyncio为何仍串行执行S3调用
你的代码看似使用了asyncio的协程与asyncio.gather来实现并发,但实际所有S3请求还是串行执行,核心原因如下:
核心问题:boto3客户端是阻塞式的
boto3的client方法(比如list_objects_v2)都是同步阻塞调用,当你在异步函数_nb_of_objects里直接执行self.s3.list_objects_v2时,这个调用会占用当前线程直到请求完全完成,asyncio的事件循环根本没有机会切换到其他协程,最终导致所有任务串行执行。
asyncio的并发依赖于非阻塞IO:只有当协程遇到await一个真正的异步操作(不会阻塞线程的操作)时,事件循环才会暂停当前协程,转而执行其他就绪的协程。而boto3的同步调用完全不符合这个条件。
关于“无需aioboto3也能异步”的方案说明
那些教程提到的方案,本质是将阻塞的boto3调用放到线程池里执行,通过线程来隔离阻塞操作,让asyncio事件循环可以正常调度协程。Python 3.9+提供了asyncio.to_thread简化这个操作,低版本则可以用run_in_executor。
修改后的代码示例(Python 3.9+)
import asyncio, boto3 class wtf: def __init__(self): self.s3 = boto3.client("s3") print(asyncio.run(self.get_all([f"key{n}" for n in range(10)]))) async def get_all(self, keys: list[str]) -> list[int]: print("Start get_all") res = await asyncio.gather(*[self.get_stuff(key) for key in keys]) print("End get_all") return res async def get_stuff(self, key) -> int: print(f" Start get_stuff {key}") res = await self._nb_of_objects(key) print(f" End get_stuff {key}") return res async def _nb_of_objects(self, key) -> int: print(f" Start nb obj {key}") # 将阻塞的boto3调用移到线程池执行 s3_objects = await asyncio.to_thread( self.s3.list_objects_v2, Bucket="BUCKET_NAME", Prefix=key ) l = len(s3_objects.get("Contents", [])) print(f" End nb obj {key}") return l if __name__ == "__main__": wtf()
低版本Python兼容方案
如果使用Python 3.9以下版本,替换_nb_of_objects为:
async def _nb_of_objects(self, key) -> int: print(f" Start nb obj {key}") loop = asyncio.get_running_loop() s3_objects = await loop.run_in_executor( None, # 使用默认线程池 self.s3.list_objects_v2, Bucket="BUCKET_NAME", Prefix=key ) l = len(s3_objects.get("Contents", [])) print(f" End nb obj {key}") return l
关于aioboto3的选择
aioboto3是官方维护的异步AWS SDK,内部实现了真正的非阻塞IO,不需要依赖线程池,在高并发场景下性能更优,是异步操作AWS资源的推荐方案。而线程池包装boto3的方式,属于“兼容式异步”,适合不想引入额外依赖的简单场景。
内容的提问来源于stack exchange,提问作者Guillaume
相关产品推荐
相关产品推荐

