Python异步IO读取硬盘文件比线程慢的原因及大量任务使用疑问
异步IO与线程池处理百万图片插入Redis的性能问题
问题背景
我有大约100万张图片,需要用Python读取并将字节数据插入Redis。这是纯IO任务,我尝试了两种方案:线程池和asyncio,但发现线程池方法比asyncio快10倍以上。示例代码如下:
import pickle import os import os.path as osp import re import redis import asyncio from multiprocessing.dummy import Pool r = redis.StrictRedis(host='localhost', port=6379, db=1) data_root = './datasets/images' print('obtain name and paths') paths_names = [] for root, dis, fls in os.walk(data_root): for fl in fls: if re.search('JPEG$', fl) is None: continue pth = osp.join(root, fl) name = re.sub('\.+/', '', pth) name = re.sub('/', '-', name) name = 'redis-' + name paths_names.append((pth, name)) print('num samples in total: ', len(paths_names)) ### 异步方案(更慢) print('insert into redis') async def insert_one(path_name): pth, name = path_name if r.get(name): return with open(pth, 'rb') as fr: binary = fr.read() r.set(name, binary) async def func(cid, n_co): num = len(paths_names) for i in range(cid, num, n_co): await insert_one(paths_names[i]) n_co = 256 loop = asyncio.get_event_loop() tasks = [loop.create_task(func(cid, n_co)) for cid in range(n_co)] fut = asyncio.gather(*tasks) loop.run_until_complete(fut) loop.close() ### 线程池方案(快10倍以上) def insert_one(path_name): pth, name = path_name if r.get(name): return with open(pth, 'rb') as fr: binary = fr.read() r.set(name, binary) def func(cid, n_co): num = len(paths_names) for i in range(cid, num, n_co): insert_one(paths_names[i]) with Pool(128) as pool: pool.map(func, paths_names)
我有两个困惑:
- 异步IO方案存在什么问题,导致其比线程方案慢?
- 是否建议向
gather函数添加数百万甚至更多任务?例如:
num_parallel = 1000000000 tasks = [loop.create_task(fetch_func(cid, num_parallel)) for cid in range(num_parallel)] await asyncio.gather(*tasks)
问题解答
1. 异步IO方案慢的核心原因
你的异步代码完全没有实现真正的异步IO:
- 使用的
redis.StrictRedis是同步阻塞客户端,r.get()和r.set()会直接卡住整个事件循环,asyncio的单线程调度优势完全无法发挥——事件循环被同步操作占满,根本没法切换到其他任务,本质和单线程同步执行没区别。 - 同步文件操作
open()和fr.read()同样会阻塞事件循环,进一步拖慢执行速度。 - 线程池方案中,每个线程的阻塞IO会被操作系统调度到其他线程执行,多个线程可以同时等待IO完成,真正实现了并行处理。
要让asyncio方案高效运行,必须替换为全异步栈:
- 使用异步Redis客户端(如
aredis或redis-py的异步版本) - 用
aiofiles替代同步文件操作 - 确保所有IO调用都是非阻塞的异步接口,让事件循环能在等待IO时切换任务。
2. 绝对不建议向gather添加百万级任务
直接创建数百万甚至更多asyncio任务会引发严重问题:
- 内存耗尽:每个任务都需要占用内存资源,百万级任务会快速吃光系统内存,导致程序崩溃或系统卡顿。
- 调度开销剧增:事件循环需要维护和调度大量任务,调度的时间成本会远超过任务执行的实际开销,反而大幅降低整体效率。
合理的做法是:
- 使用信号量/任务池限制并发任务数量(比如设置几十到几百的并发数),避免同时运行过多任务。
- 分批次处理任务,或用异步迭代器逐步生成任务,控制并发量在系统可承受范围内。
内容的提问来源于stack exchange,提问作者coin cheung
相关产品推荐
相关产品推荐

