You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何使用asyncio.gather处理pandas DataFrame行无法并发执行?

为什么aiohttp异步请求未并发执行,耗时与同步一致?

核心原因

你的异步代码未实现并发的关键问题在于aiohttp默认的TCP连接器限制了并发连接数。aiohttp.ClientSession默认使用的TCPConnector的limit参数值为10,意味着同一时间最多只能建立10个HTTP连接。你的测试用例是100个请求,会被分成10批串行执行,每批10个请求耗时约0.1秒,总耗时刚好10秒,和同步循环表现一致。

修复方案

创建ClientSession时,显式指定TCPConnector并调整limit(全局并发数)或limit_per_host(单域名并发数)参数,放开并发限制:

import asyncio
import aiohttp
import pandas as pd
import time

# 模拟100行数据,每行请求延迟0.1秒的端点
df = pd.DataFrame({
    'id': range(100),
    'url': ['http://httpbin.org/delay/0.1'] * 100
})

async def fetch(session, url, row_id):
    async with session.get(url) as response:
        await response.text()
        return row_id

async def process_dataframe(df):
    # 显式配置连接器,设置全局最大并发连接数
    connector = aiohttp.TCPConnector(limit=100)
    # 若仅针对单域名请求,也可使用limit_per_host精准控制单域名并发
    # connector = aiohttp.TCPConnector(limit_per_host=100)
    async with aiohttp.ClientSession(connector=connector) as session:
        tasks = [asyncio.create_task(fetch(session, row['url'], row['id'])) 
                 for _, row in df.iterrows()]
        results = await asyncio.gather(*tasks)
    return results

start = time.time()
results = asyncio.run(process_dataframe(df))
print(f"Async time: {time.time() - start:.2f} seconds")

关键细节说明

  1. 并发数调整原则:

    • 测试用例中设置limit=100即可实现全并发,耗时会降到0.1-0.2秒区间,符合预期。
    • 针对50万行的大规模场景,不建议直接将并发数设为50万,会导致内存占用飙升、触发服务器限流甚至被封禁。建议根据目标服务器承受能力,设置合理的并发数(如200-500),同时将DataFrame拆分批次处理,避免一次性创建过多异步任务。
  2. 参数区别:

    • limit:控制所有域名的总并发连接数。
    • limit_per_host:控制单个域名的并发连接数,更适合多域名请求场景,避免对单一服务器造成过载压力。

验证结果

修改后运行脚本,100个延迟请求会同时发起,总耗时稳定在0.1-0.2秒,符合异步并发的预期效果。

内容的提问来源于stack exchange,提问作者Джон Сноу

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.01 13:14:54