You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+Playwright结合并发与并行执行数千请求?

关于AsyncIO、aiomultiprocess结合Playwright的爬虫优化问题

测试场景说明

异步执行100次请求并打印网页标题,对比两种实现的耗时:

AsyncIO实现

import asyncio
import time
from playwright.async_api import async_playwright, Browser

async def worker(browser: Browser, i: int):
    context = await browser.new_context()
    page = await context.new_page()
    await page.goto("https://www.google.com/")
    print(await page.title())
    await context.close()

async def main():
    async with async_playwright() as playwright:
        browser = await playwright.chromium.launch(channel="chrome")
        await asyncio.wait(
            [asyncio.create_task(worker(browser, i)) for i in range(100)],
            return_when=asyncio.ALL_COMPLETED,
        )
        await browser.close()

start = time.perf_counter()
asyncio.run(main())
end = time.perf_counter()
print(end - start)

耗时:8.336619800000335秒

Aiomultiprocess实现(原版本)

import asyncio
import time
from playwright.async_api import async_playwright
from aiomultiprocess import Pool

async def run(playwright, url):
    chromium = playwright.chromium # or "firefox" or "webkit".
    browser = await chromium.launch(channel="chrome")
    page = await browser.new_page()
    await page.goto(url)
   
    print(await page.title())
    await browser.close()

async def mains(url):
    async with async_playwright() as playwright:
        return await run(playwright,url)

async def main():
    urls = ["https://www.google.com/" for i in range(100)]
    async with Pool() as pool:
        async for result in pool.map(mains, urls):
            pass

if __name__ == '__main__':
    start = time.perf_counter()
    asyncio.run(main())
    end = time.perf_counter()
    print(end - start, end="")

耗时:17.115715199999613秒


问题解答

1. 当前实现是否正确?若不正确,如何优化以提升速度?

  • AsyncIO实现:核心逻辑正确,仅存在冗余import asyncio的小问题。该实现复用单个浏览器实例,通过创建多个上下文处理请求,是Playwright异步场景的高效写法,耗时表现合理。
  • Aiomultiprocess原实现:逻辑存在严重效率问题——每个请求都启动并关闭一个完整的浏览器实例,浏览器启动/销毁的开销极大,直接导致耗时翻倍。

优化后的Aiomultiprocess实现:每个进程仅初始化一次浏览器,进程内复用浏览器处理多个请求,大幅降低资源开销:

import asyncio
import time
from playwright.async_api import async_playwright, Browser
from aiomultiprocess import Pool

# 每个进程初始化一次浏览器
browser: Browser = None

async def init_browser():
    global browser
    playwright = await async_playwright().start()
    browser = await playwright.chromium.launch(channel="chrome")

async def worker(url):
    global browser
    context = await browser.new_context()
    page = await context.new_page()
    await page.goto(url)
    print(await page.title())
    await context.close()

async def main():
    urls = ["https://www.google.com/" for i in range(100)]
    # 初始化每个进程的浏览器
    async with Pool(initializer=init_browser) as pool:
        async for _ in pool.map(worker, urls):
            pass
    # 关闭浏览器(需确保进程退出前执行)
    await browser.close()

if __name__ == '__main__':
    start = time.perf_counter()
    asyncio.run(main())
    end = time.perf_counter()
    print(end - start)

优化后,多进程方案的耗时会大幅接近甚至超过纯AsyncIO方案(取决于CPU核心数和单进程并发数)。

2. 是否应仅使用AsyncIO?

需根据场景判断:

  • 若仅处理纯IO密集型任务(如网络请求、页面渲染),纯AsyncIO完全足够。Playwright本身是异步设计,复用单浏览器实例的效率极高,多进程带来的进程通信、资源初始化开销反而会拖慢速度,当前测试场景就是如此。
  • 若存在CPU密集型任务(如大文件解析、复杂数据计算),或依赖无法异步的同步阻塞代码,此时结合多进程+AsyncIO才能充分利用多核CPU资源,提升整体效率。

3. 若扩展至数千次请求,如何用Playwright或其他爬虫库实现?

基于Playwright的实现方案

  1. 纯AsyncIO方案:

    • 用asyncio.Semaphore限制并发数(避免触发网站反爬或耗尽本地资源),比如限制同时执行30个请求;
    • 复用单个浏览器实例,批量创建上下文/页面,任务完成后统一清理资源;
    • 核心逻辑示例:
      async def main():
          semaphore = asyncio.Semaphore(30)
          async with async_playwright() as playwright:
              browser = await playwright.chromium.launch(channel="chrome")
              tasks = []
              for url in urls:
                  task = asyncio.create_task(worker_with_semaphore(browser, url, semaphore))
                  tasks.append(task)
              await asyncio.gather(*tasks)
              await browser.close()
      
      async def worker_with_semaphore(browser, url, semaphore):
          async with semaphore:
              context = await browser.new_context()
              page = await context.new_page()
              await page.goto(url)
              print(await page.title())
              await context.close()
      
  2. 多进程+AsyncIO混合方案:

    • 每个进程启动一个浏览器,进程内用AsyncIO+Semaphore控制单进程并发数;
    • 总并发数=进程数×单进程并发数,比如4个进程,每个进程并发30,总并发120;
    • 适合CPU资源充足、需要更高并发的场景,同时避免单进程异步的GIL限制。

基于其他爬虫库的实现

如果不需要页面渲染(仅需HTTP请求),推荐使用aiohttp(轻量异步HTTP库),资源开销远低于Playwright:

  • 用aiohttp.ClientSession复用连接池,结合asyncio.Semaphore控制并发;
  • 若需更高并发,可结合aiomultiprocess,每个进程维护一个ClientSession,处理部分请求。

内容的提问来源于stack exchange,提问作者AIboi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 16:40:56