You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Python 3.11+实现高并发HTTP请求(每秒150-500+)?

高并发HTTPS请求解决方案(Python 3.11+)

方案1:线程池 + Requests

无需复杂配置,快速实现并发请求,适合大多数场景。

代码实现

import requests
from concurrent.futures import ThreadPoolExecutor, as_completed
from typing import List, str

# 复用会话减少TCP连接开销
session = requests.Session()

def fetch_html(url: str) -> str:
    try:
        response = session.get(url, timeout=10)
        response.raise_for_status()
        return response.text
    except Exception as e:
        print(f"请求{url}失败: {str(e)}")
        return ""

def batch_fetch(urls: List[str], max_workers: int = 80) -> List[str]:
    html_results = []
    with ThreadPoolExecutor(max_workers=max_workers) as executor:
        future_map = {executor.submit(fetch_html, url): url for url in urls}
        for future in as_completed(future_map):
            html_results.append(future.result())
    return html_results

# 使用示例
if __name__ == "__main__":
    target_urls = [
        "https://example.com",
        "https://google.com"
        # 添加更多URL
    ]
    results = batch_fetch(target_urls, max_workers=80)
    print(f"完成{len(results)}个请求")

优化提示

  • 调整max_workers:根据硬件性能和目标网站反爬策略调整,50-100是常规合理范围
  • 增加重试机制:可借助tenacity库实现失败自动重试,降低偶发错误影响
  • 自定义请求头:添加User-Agent等字段,模拟浏览器请求规避反爬

方案2:Httpx异步客户端

比AIOHTTP配置更简洁,异步IO更节省资源,支持更高并发(轻松达到500+请求/秒)

代码实现

import httpx
import asyncio
from typing import List, str

async def fetch_html_async(url: str, client: httpx.AsyncClient) -> str:
    try:
        response = await client.get(url, timeout=10)
        response.raise_for_status()
        return response.text
    except Exception as e:
        print(f"请求{url}失败: {str(e)}")
        return ""

async def batch_fetch_async(urls: List[str], conn_limit: int = 100) -> List[str]:
    async with httpx.AsyncClient(limits=httpx.Limits(max_connections=conn_limit)) as client:
        tasks = [fetch_html_async(url, client) for url in urls]
        results = await asyncio.gather(*tasks)
    return results

# 使用示例
if __name__ == "__main__":
    target_urls = [
        "https://example.com",
        "https://google.com"
        # 添加更多URL
    ]
    results = asyncio.run(batch_fetch_async(target_urls, conn_limit=100))
    print(f"完成{len(results)}个请求")

核心优势

  • 内置连接池与超时管理,无需额外配置
  • 支持HTTP/2协议,进一步提升请求效率
  • 异步模型比线程池占用更少CPU和内存,适合超高并发场景

通用注意事项

  • 反爬规避:高并发请求易触发IP封禁,建议搭配代理IP池使用
  • 资源控制:并发数过高会耗尽本地网络或硬件资源,需根据实际情况调整参数
  • 异常处理:必须捕获请求异常,避免单个请求失败导致整个任务崩溃

内容的提问来源于stack exchange,提问作者Pengalor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 18:40:25