You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python asyncio中to_thread为何比create_task更快?多实现疑问解析

问题描述

我有一个调用REST接口并返回结果的阻塞函数,模拟代码如下:

import random, time, asyncio
def test_func(x):
    time.sleep(2*random.random())
    return x

我尝试使用Python asyncio 将该阻塞调用转为异步非阻塞调用,采用了两种实现方式:

1. 第一种实现

class Output:
    output = None

def test_func_thread(x, out: Output):
    out.output = test_func(x)

async def imp1():
    coroutines = []
    scores = []
    for _ in range(10):
        scores.append(Output())
        coroutines.append(asyncio.to_thread(test_func_thread, _, scores[-1]))
    for coroutine in asyncio.as_completed(coroutines):
        await coroutine
    return [_.output for _ in scores]

2. 第二种实现

async def async_score(x):
    return test_func(x)

async def imp2():
    tasks = []
    results = []
    for _ in range(10):
        tasks.append(asyncio.create_task(async_score(_)))

    for t in tasks:
        if not t.done():
            await t
        results.append(t.result())
    return results 

同步阻塞版本

def imp3():
    result = []
    for _ in range(10):
        result.append(test_func(_))
    return result

各版本运行耗时如下:

print(asyncio.run(imp1()))
> [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]
%timeit asyncio.run(imp1())
> 1 loops, best of 5: 1.45 s per loop

print(asyncio.run(imp2()))
> [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]
%timeit asyncio.run(imp2())
> 1 loops, best of 5: 7.11 s per loop

print(imp3())
> [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]
%timeit imp3()
> 1 loops, best of 5: 8.66 s per loop

我原本预期imp2()与imp1()性能相近,但imp2()实际更接近同步阻塞版本,使用asyncio几乎无性能提升。我对asyncio并不熟悉,因此有以下疑问:

  1. 为何imp2()表现得像同步阻塞版本?
  2. 此类场景下的Pythonic实现方式是什么?
  3. 如果imp2()更符合Python风格,如何使其性能达到imp1()的水平?
  4. 对于imp1(),我不得不使用Output类的技巧来保持结果与调用顺序一致,这似乎不够Pythonic,有没有更好的方法?

更新:

对于imp2,我还尝试了以下实现:

async def imp2():
    tasks = [async_score(_) for _ in range(10)]
    results = await asyncio.gather(*tasks)
    return results

但它仍然表现得像同步阻塞调用,没有任何性能提升。


问题解答

1. 为何imp2()表现得像同步阻塞版本?

因为async_score虽然被定义为async函数,但内部直接调用了阻塞的同步函数test_func,且未将阻塞操作放到线程或进程中执行。

asyncio的异步机制依赖于可等待对象主动让出CPU,但test_func里的time.sleep是阻塞整个线程的同步操作——一旦执行到test_func,asyncio的事件循环会被完全卡住,直到这个阻塞操作结束才能处理下一个任务。不管用create_task还是asyncio.gather,本质都是在同一个线程里依次执行这些阻塞调用,和同步版本的区别仅多了asyncio的任务调度开销,所以性能和同步版本几乎一致。

2. 此类场景下的Pythonic实现方式是什么?

处理阻塞IO操作(比如REST接口调用、文件读写)的Pythonic异步方式,优先使用支持异步的专用库(如调用REST接口用aiohttp);如果必须基于现有阻塞函数改造,最简洁的方式是结合asyncio.to_thread和asyncio.gather,既保证并发执行,又能直接按调用顺序获取结果。

3. 如果imp2()更符合Python风格,如何使其性能达到imp1()的水平?

核心是把阻塞的test_func放到线程中执行,避免卡住asyncio事件循环。只需修改async_score函数:

async def async_score(x):
    # 用asyncio.to_thread将阻塞函数放到线程池运行
    return await asyncio.to_thread(test_func, x)

修改后,不管是原imp2实现还是更新后的asyncio.gather版本,都会和imp1一样实现并发执行,性能大幅提升。

4. 对于imp1(),有没有更Pythonic的方法保持结果顺序?

有两种更简洁的方式:

  • 直接用asyncio.gather配合asyncio.to_thread,天然按调用顺序返回结果:
async def imp1_improved():
    tasks = [asyncio.to_thread(test_func, x) for x in range(10)]
    return await asyncio.gather(*tasks)
  • 如果需要用asyncio.as_completed处理完成的任务,可通过字典记录任务与对应索引,再填充结果列表:
async def imp1_improved_as_completed():
    tasks = {asyncio.to_thread(test_func, x): x for x in range(10)}
    results = [0]*10
    for coro in asyncio.as_completed(tasks):
        res = await coro
        idx = tasks[coro]
        results[idx] = res
    return results

这两种方式都比自定义Output类更符合Python惯用写法。

内容的提问来源于stack exchange,提问作者qkhhly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 18:43:25