You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

异步脚本批量请求不同域名URL时出现异常超时问题求助

批量异步请求后期超时问题的解决方案

问题描述

我编写了一个异步脚本,向223个不同域名的URL各发送1次请求。大约发送150次请求后,部分URL开始出现超时,但单独处理这些超时URL时却能正常完成请求。尝试过使用aiohttp替代httpx、添加asyncio.sleep()、为每个请求创建单独客户端,均未解决问题。

相关代码:

请求函数

async def fetch(url, client):
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/97.0.4692.99 Safari/537.36",  
        "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7",
    }

    response = await client.get(url, follow_redirects=True, headers=headers)
    return response

运行代码

urls_list = [site.strip() for site in start_urls.replace(',', ' ').split() if site.strip()]
if second_input:
    pattern_list = second_input
else:
    pattern_list = []
           
async with httpx.AsyncClient() as client:
    tasks = [process_url(url, skip_empty_results, depth, pattern_list, dataset, 5, client) for url in urls_list]
    await asyncio.gather(*tasks)

其中fetch()函数在process_url()中被调用。

问题原因分析

这种批量异步请求后期超时的核心原因通常是系统资源瓶颈或DNS解析过载:

  • 文件句柄限制:每个HTTP连接会占用一个系统文件句柄,当并发请求数超过系统允许的最大打开文件数时,新连接会被阻塞直至超时。
  • DNS解析瓶颈:大量并发请求同时触发DNS查询,本地解析服务或DNS缓存无法及时处理,导致解析超时进而引发请求超时。
  • 无限制并发的资源消耗:一次性发起所有请求会瞬间耗尽系统套接字、内存等资源,导致后续请求无法正常建立连接。

解决方案

1. 限制并发请求数量

使用asyncio.Semaphore控制同时运行的任务数,避免一次性发起全部请求,降低系统资源压力:

async def process_url_with_semaphore(url, semaphore, *args, **kwargs):
    async with semaphore:
        return await process_url(url, *args, **kwargs)

# 修改后的运行代码
async def main():
    urls_list = [site.strip() for site in start_urls.replace(',', ' ').split() if site.strip()]
    pattern_list = second_input if second_input else []
           
    async with httpx.AsyncClient() as client:
        # 限制同时并发20个请求,可根据系统配置调整(建议10-30之间)
        semaphore = asyncio.Semaphore(20)
        tasks = [process_url_with_semaphore(url, semaphore, skip_empty_results, depth, pattern_list, dataset, 5, client) for url in urls_list]
        await asyncio.gather(*tasks)

2. 优化DNS解析效率

通过自定义DNS解析器提升批量解析速度,避免请求时才触发解析:
首先安装依赖:

pip install aiodns

然后修改客户端初始化逻辑:

import aiodns

async def create_custom_resolver():
    resolver = aiodns.DNSResolver()
    # 指定公共DNS服务器,加快解析速度
    resolver.nameservers = ['8.8.8.8', '1.1.1.1']
    return resolver

async def main():
    urls_list = [site.strip() for site in start_urls.replace(',', ' ').split() if site.strip()]
    pattern_list = second_input if second_input else []
    
    resolver = await create_custom_resolver()
    async with httpx.AsyncClient(resolver=resolver) as client:
        semaphore = asyncio.Semaphore(20)
        tasks = [process_url_with_semaphore(url, semaphore, skip_empty_results, depth, pattern_list, dataset, 5, client) for url in urls_list]
        await asyncio.gather(*tasks)

3. 添加请求超时与重试机制

针对个别请求的超时情况,添加重试逻辑提高成功率:
首先安装重试库:

pip install tenacity

修改fetch函数:

from tenacity import retry, stop_after_attempt, wait_exponential

@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10))
async def fetch(url, client):
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/97.0.4692.99 Safari/537.36",  
        "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7",
    }
    try:
        # 设置10秒超时时间
        response = await client.get(url, follow_redirects=True, headers=headers, timeout=10.0)
        return response
    except httpx.TimeoutException:
        raise  # 抛出异常触发重试

4. 调整系统资源限制(Linux)

如果上述方案仍无法解决,可临时调高系统文件句柄限制:

# 当前终端会话生效
ulimit -n 4096

若需永久生效,编辑/etc/security/limits.conf,添加以下内容:

* soft nofile 4096
* hard nofile 8192

内容的提问来源于stack exchange,提问作者Kirada

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 15:07:33