You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Playwright连接外部浏览器并行检查书签URL有效性

问题

需要用Playwright验证最多1000个书签URL的有效性,要求:

  • 仅创建一次浏览器实例
  • 3-8个并发检查
  • 高效处理大量URL

当前使用基于browserless的外部Firefox实例,尝试用多个浏览器上下文处理URL验证,但运行代码时始终报错:Target page, context or browser has been closed

原型代码如下:

async def check_bookmark(context, url):
    page = await context.new_page()
    try:
        await page.goto(url)
        status = await page.evaluate('() => document.readyState')
        if status == 'complete':
            print(f"{url} is valid")
        else:
            print(f"{url} is invalid")
    except Exception as e:
        print(f"{url} is invalid: {str(e)}")
    finally:
        await page.close()

async def main():
     bookmark_urls = [
        "https://google.com",
        "https://youtube.com",
        "https://facebook.com",
        "https://instagram.com",
        "https://baidu.com",
        "https://wikipedia.org",
        "https://twitter.com",
        "https://yahoo.com",
        "https://yandex.ru",
        "https://whatsapp.com",
        "https://x.com",
        "https://chatgpt.com",
        "https://tiktok.com",
        "https://reddit.com",
        "https://amazon.com",
        "https://yahoo.co.jp",
        "https://live.com",
        "https://linkedin.com/"
    ]

    async with async_playwright() as p:
        browser = await p.firefox.connect('ws://localhost:3000/playwright/firefox')
        
        contexts = []

        for i in range(0, 3):
            context = await browser.new_context()
            contexts.append(context)

        tasks = []
        num_contexts = len(contexts)
        num_urls = len(bookmark_urls)
        urls_per_context = num_urls // num_contexts

        for i in range(num_contexts):
            start_index = i * urls_per_context
            end_index = start_index + urls_per_context
            urls_subset = bookmark_urls[start_index:end_index]

            for url in urls_subset:
                tasks.append(check_bookmark(contexts[i], url))

        await asyncio.gather(*tasks)

        await browser.close()

if __name__ == '__main__':
    asyncio.run(main())
解决方案

问题原因

  1. 并发过载:单个浏览器上下文同时处理多个页面任务,超出了browserless或Firefox的资源承载上限,导致上下文被强制关闭。
  2. 任务分配逻辑缺陷:将多个URL分配给同一个上下文并行处理,忽略了单上下文的并发页面限制。
  3. 资源未及时释放:未在任务完成后主动关闭上下文,引发服务端资源泄漏,进而触发关闭错误。

修复后的代码

采用信号量控制并发数,每个任务使用独立临时上下文,确保资源使用在合理范围内:

import asyncio
from playwright.async_api import async_playwright

CONCURRENCY_LIMIT = 5  # 控制在3-8之间,可根据browserless配置调整

async def check_bookmark(browser, url):
    # 每个任务创建独立上下文,隔离资源避免互相干扰
    async with await browser.new_context() as context:
        async with await context.new_page() as page:
            try:
                # 设置导航超时,避免任务长时间挂起
                response = await page.goto(url, timeout=10000, wait_until='domcontentloaded')
                if response and response.status == 200:
                    print(f"{url} 有效")
                    return {"url": url, "valid": True, "status": response.status}
                else:
                    status_code = response.status if response else "无响应"
                    print(f"{url} 无效,状态码: {status_code}")
                    return {"url": url, "valid": False, "status": status_code}
            except Exception as e:
                print(f"{url} 无效: {str(e)}")
                return {"url": url, "valid": False, "error": str(e)}

async def main():
    bookmark_urls = [
        "https://google.com",
        "https://youtube.com",
        "https://facebook.com",
        "https://instagram.com",
        "https://baidu.com",
        "https://wikipedia.org",
        "https://twitter.com",
        "https://yahoo.com",
        "https://yandex.ru",
        "https://whatsapp.com",
        "https://x.com",
        "https://chatgpt.com",
        "https://tiktok.com",
        "https://reddit.com",
        "https://amazon.com",
        "https://yahoo.co.jp",
        "https://live.com",
        "https://linkedin.com/"
    ]

    async with async_playwright() as p:
        browser = await p.firefox.connect('ws://localhost:3000/playwright/firefox')
        
        # 用信号量严格控制并发数
        semaphore = asyncio.Semaphore(CONCURRENCY_LIMIT)
        
        async def bounded_check(url):
            async with semaphore:
                return await check_bookmark(browser, url)
        
        # 生成所有任务并执行
        tasks = [bounded_check(url) for url in bookmark_urls]
        results = await asyncio.gather(*tasks)
        
        # 可选:将结果保存到文件
        # import json
        # with open('bookmark_validation_results.json', 'w') as f:
        #     json.dump(results, f, indent=2)
        
        await browser.close()

if __name__ == '__main__':
    asyncio.run(main())

优化说明

  • 并发控制:通过asyncio.Semaphore严格限制并发数,避免超出服务端资源配额。
  • 资源隔离:每个任务使用独立上下文,防止单个URL的导航异常影响其他任务。
  • 精准验证:直接通过page.goto的返回值获取HTTP状态码,比检查document.readyState更准确。
  • 自动清理:使用async with语法自动关闭上下文和页面,确保资源及时释放。
  • 结果结构化:返回包含URL、有效性、状态码/错误信息的字典,方便后续处理。

额外建议

  1. 根据browserless的CPU/内存配置,调整CONCURRENCY_LIMIT,建议不超过8。
  2. 处理1000个URL时,可每处理50-100个后短暂休眠(如1-2秒),降低服务端压力。
  3. 可添加重试机制,针对临时网络错误的URL进行1-2次重试。

内容的提问来源于stack exchange,提问作者user25674617

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 05:17:35