You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pyppeteer启用请求拦截后遇认证弹窗阻塞问题求助

解决方案

针对启用请求拦截后遇HTTP认证弹窗导致阻塞的问题,从修复代码错误和处理认证场景两方面解决,提供两种符合需求的方案:

方案一:正确处理认证弹窗,正常抓取HTML

关键修复点

  1. 修复请求拦截函数语法错误:原代码中request. Abort存在空格和方法名大小写错误,Pyppeteer中请求中止方法为request.abort(),错误会导致请求未被处理,引发无限阻塞。
  2. 完善弹窗处理逻辑:HTTP认证弹窗属于prompt类型,原代码仅处理alert/confirm,需补充该类型的处理(直接取消弹窗,或按需传入认证信息)。
  3. 确保所有请求被正确处理:请求拦截模式下,每个请求必须调用abort()/continue()/respond(),否则页面会一直等待。

修改后完整代码

from pyppeteer import launch
from utils.agents import get_user_agent
import asyncio

async def main():
    browser = await launch(
        headless=True,
        ignoreHTTPSErrors=True,
        acceptInsecureCerts=True,
        args=[
            "--no-sandbox",
            "--disable-gpu",
            "--ignore-certificate-errors",
            "--allow-running-insecure-content",
            "--disable-web-security",
            "--disable-setuid-sandbox",
            '--disable-popup-blocking',
            '--disable-dev-shm-usage',
            '--no-zygote'
        ]
    )

    page = await browser.newPage()
    user_agent = get_user_agent()
    await page.setUserAgent(user_agent)
    await page.setRequestInterception(True)

    url = "http://mogilitycapital.com"
    netcalls = []

    async def handle_request_redirects(request):
        # 中止非必要资源请求
        if request.resourceType in ['stylesheet', 'css', 'image', 'font']:
            await request.abort('blockedbyclient')
        else:
            await request.continue_()

    async def intercept_network_response(response):
        netcalls.append({
            "url": response.url,
            "method": response.request.method,
            "headers": response.headers,
            "status": response.status
        })
        # 检测是否为HTTP认证请求(401状态码)
        if response.status == 401:
            print(f"站点 {response.url} 需要HTTP认证")

    async def handle_dialog(dialog):
        # 处理HTTP认证弹窗(prompt类型)
        if dialog.type == 'prompt':
            await dialog.dismiss()  # 取消弹窗,部分站点会返回未授权页面
        elif dialog.type == 'alert':
            await dialog.dismiss()
        elif dialog.type == 'confirm':
            await dialog.accept()

    # 绑定事件监听
    page.on('request', lambda req: asyncio.ensure_future(handle_request_redirects(req)))
    page.on('response', lambda res: asyncio.ensure_future(intercept_network_response(res)))
    page.on('dialog', lambda dlg: asyncio.ensure_future(handle_dialog(dlg)))

    try:
        resp = await page.goto(url, timeout=60000)
        # 获取原始HTML
        html = await page.content()
        print("抓取到HTML内容")
    except Exception as e:
        print(f"抓取失败: {e}")
    finally:
        await browser.close()

if __name__ == "__main__":
    asyncio.run(main())

方案二:拦截带HTTP认证的网站,直接跳过

针对百万级域名的批量抓取需求,可在响应阶段检测401状态码(HTTP未授权),直接标记并跳过该站点的后续处理,避免资源浪费:

核心逻辑

  1. 在响应拦截函数中检查状态码,若为401则记录该站点为需要认证的站点。
  2. 捕获goto时的超时或认证错误,直接关闭页面并跳过后续处理。

关键代码片段

async def intercept_network_response(response):
    netcalls.append({
        "url": response.url,
        "method": response.request.method,
        "headers": response.headers,
        "status": response.status
    })
    # 检测401状态码,标记站点
    if response.status == 401:
        print(f"拦截需认证站点: {response.url}")
        # 可将域名存入黑名单,后续批量处理时跳过
        with open('auth_sites.txt', 'a') as f:
            f.write(f"{response.url}\n")

# 在goto时增加错误捕获
try:
    resp = await page.goto(url, timeout=60000)
    if resp.status == 401:
        print(f"跳过需认证站点: {url}")
    else:
        html = await page.content()
        # 处理HTML和网络请求数据
except asyncio.TimeoutError:
    print(f"站点 {url} 超时,跳过")
except Exception as e:
    print(f"站点 {url} 处理失败: {e}")

内容的提问来源于stack exchange,提问作者Pyd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 15:09:51