Python asyncio中to_thread为何比create_task更快?多实现疑问解析
我有一个调用REST接口并返回结果的阻塞函数,模拟代码如下:
import random, time, asyncio def test_func(x): time.sleep(2*random.random()) return x
我尝试使用Python asyncio 将该阻塞调用转为异步非阻塞调用,采用了两种实现方式:
1. 第一种实现
class Output: output = None def test_func_thread(x, out: Output): out.output = test_func(x) async def imp1(): coroutines = [] scores = [] for _ in range(10): scores.append(Output()) coroutines.append(asyncio.to_thread(test_func_thread, _, scores[-1])) for coroutine in asyncio.as_completed(coroutines): await coroutine return [_.output for _ in scores]
2. 第二种实现
async def async_score(x): return test_func(x) async def imp2(): tasks = [] results = [] for _ in range(10): tasks.append(asyncio.create_task(async_score(_))) for t in tasks: if not t.done(): await t results.append(t.result()) return results
同步阻塞版本
def imp3(): result = [] for _ in range(10): result.append(test_func(_)) return result
各版本运行耗时如下:
print(asyncio.run(imp1())) > [0, 1, 2, 3, 4, 5, 6, 7, 8, 9] %timeit asyncio.run(imp1()) > 1 loops, best of 5: 1.45 s per loop print(asyncio.run(imp2())) > [0, 1, 2, 3, 4, 5, 6, 7, 8, 9] %timeit asyncio.run(imp2()) > 1 loops, best of 5: 7.11 s per loop print(imp3()) > [0, 1, 2, 3, 4, 5, 6, 7, 8, 9] %timeit imp3() > 1 loops, best of 5: 8.66 s per loop
我原本预期imp2()与imp1()性能相近,但imp2()实际更接近同步阻塞版本,使用asyncio几乎无性能提升。我对asyncio并不熟悉,因此有以下疑问:
- 为何
imp2()表现得像同步阻塞版本? - 此类场景下的Pythonic实现方式是什么?
- 如果
imp2()更符合Python风格,如何使其性能达到imp1()的水平? - 对于
imp1(),我不得不使用Output类的技巧来保持结果与调用顺序一致,这似乎不够Pythonic,有没有更好的方法?
更新:
对于imp2,我还尝试了以下实现:
async def imp2(): tasks = [async_score(_) for _ in range(10)] results = await asyncio.gather(*tasks) return results
但它仍然表现得像同步阻塞调用,没有任何性能提升。
1. 为何imp2()表现得像同步阻塞版本?
因为async_score虽然被定义为async函数,但内部直接调用了阻塞的同步函数test_func,且未将阻塞操作放到线程或进程中执行。
asyncio的异步机制依赖于可等待对象主动让出CPU,但test_func里的time.sleep是阻塞整个线程的同步操作——一旦执行到test_func,asyncio的事件循环会被完全卡住,直到这个阻塞操作结束才能处理下一个任务。不管用create_task还是asyncio.gather,本质都是在同一个线程里依次执行这些阻塞调用,和同步版本的区别仅多了asyncio的任务调度开销,所以性能和同步版本几乎一致。
2. 此类场景下的Pythonic实现方式是什么?
处理阻塞IO操作(比如REST接口调用、文件读写)的Pythonic异步方式,优先使用支持异步的专用库(如调用REST接口用aiohttp);如果必须基于现有阻塞函数改造,最简洁的方式是结合asyncio.to_thread和asyncio.gather,既保证并发执行,又能直接按调用顺序获取结果。
3. 如果imp2()更符合Python风格,如何使其性能达到imp1()的水平?
核心是把阻塞的test_func放到线程中执行,避免卡住asyncio事件循环。只需修改async_score函数:
async def async_score(x): # 用asyncio.to_thread将阻塞函数放到线程池运行 return await asyncio.to_thread(test_func, x)
修改后,不管是原imp2实现还是更新后的asyncio.gather版本,都会和imp1一样实现并发执行,性能大幅提升。
4. 对于imp1(),有没有更Pythonic的方法保持结果顺序?
有两种更简洁的方式:
- 直接用
asyncio.gather配合asyncio.to_thread,天然按调用顺序返回结果:
async def imp1_improved(): tasks = [asyncio.to_thread(test_func, x) for x in range(10)] return await asyncio.gather(*tasks)
- 如果需要用
asyncio.as_completed处理完成的任务,可通过字典记录任务与对应索引,再填充结果列表:
async def imp1_improved_as_completed(): tasks = {asyncio.to_thread(test_func, x): x for x in range(10)} results = [0]*10 for coro in asyncio.as_completed(tasks): res = await coro idx = tasks[coro] results[idx] = res return results
这两种方式都比自定义Output类更符合Python惯用写法。
内容的提问来源于stack exchange,提问作者qkhhly

