You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pyppeteer加载页面后HTML空白,元素选择超时问题排查

问题

我正在使用Pyppeteer(Puppeteer的非官方Python移植版本)进行网页抓取,尝试选择元素时遇到问题。例如,我执行elements = await page.querySelectorAll('.tab')等待class为"tab"的元素,但始终返回空且触发超时错误。实际上,等待该网站任意class的元素都会超时。我尝试导出页面HTML排查,发现不仅目标元素缺失,打开HTML文件时页面完全空白。以下是我的代码片段,请问代码执行逻辑是否存在缺陷?

import asyncio
from pyppeteer import launch
import requests
import sys
import json

print("Test")
print("Starting script...")

def print_cookies_as_json(cookies):
    cookies_json = json.dumps(cookies, indent=4)
    print("Cookies in JSON format:")
    print(cookies_json)

async def main():
    try:
        print("Launching browser...")
        browser = await launch(headless=False)
        page = await browser.newPage()
    except Exception as e:
        print(f"Error launching browser or creating new page: {e}")
        return

    try:
        print("Reading cookies from file...")
        # Load cookies from JSON file
        try:
            with open('/path/to/usrCookies.json', 'r') as f:
                cookies = json.load(f)
            if not cookies:
                raise ValueError("No valid cookies found in the file.")
            await page.setCookie(*cookies)
        except (FileNotFoundError, ValueError) as e:
            print(f"Error reading cookies: {e}")
            return

        # Print cookies in JSON format
        print_cookies_as_json(cookies)

        print("Navigating to URL...")
        url = 'https://example.com'
        await page.goto(url, {'waitUntil': 'networkidle0'})
        response = requests.get(url, cookies={c['name']: c['value'] for c in cookies})
        json_response = response.json()
        print(json_response)

        print("Processing JSON data...")

        await page.screenshot({'path': 'screenshot_TEST.png', 'fullPage': True})
        print("Waiting for page to load...")
        await page.waitForSelector('body', {'timeout': 10000})  # wait for the body to load

        await asyncio.sleep(1)

        # Get the HTML content of the page
        print("Getting HTML content...")
        html = await page.content()

        # Write the HTML content to a file
        with open("index.html", "w", encoding="utf-8") as file:
           file.write(html)

    except Exception as e:
        print("An error occurred:", e)
        await page.screenshot({'path': 'screenshot_ERROR.png', 'fullPage': True})
        print("Screenshot saved as screenshot.png")
    finally:
        # Close the browser
        try:
            print("Closing browser...")
            await browser.close()
        except Exception as e:
            print(f"Error closing browser: {e}")

asyncio.get_event_loop().run_until_complete(main())
分析与解决方案
  • 多余的requests请求破坏会话:你在page.goto之后调用了requests.get,这会创建一个独立于Pyppeteer浏览器的新HTTP会话,不仅完全没用,还可能触发网站的反爬机制,导致浏览器中的会话失效,页面被清空。直接删掉这部分代码即可,Pyppeteer已经通过page.setCookie配置了会话Cookie。

  • Cookie导入可能不完整:Pyppeteer的page.setCookie要求每个Cookie必须包含domain字段(比如".example.com"),如果你的usrCookies.json里没有这个字段,Cookie设置会失效,导致访问网站时没有正确的身份验证,返回空白页面。检查Cookie文件,确保每个Cookie对象都有domain、name、value这些必填项。

  • 等待逻辑顺序混乱:page.goto已经设置了waitUntil: 'networkidle0',之后再等待body元素完全多余。而且如果中间插入了干扰操作(比如那个requests请求),页面状态已经异常,此时等待body毫无意义。应该直接在page.goto之后等待目标元素,比如await page.waitForSelector('.tab', {'timeout': 15000})。

  • 固定sleep不可靠:用asyncio.sleep(1)等待页面加载是碰运气的做法,换成等待目标元素的方式,能确保元素加载完成后再执行后续操作。

修改后的核心代码示例:

print("Navigating to URL...")
url = 'https://example.com'
# 确保Cookie包含domain字段,否则setCookie不生效
await page.setCookie(*cookies)
# networkidle2比networkidle0更适合大多数场景,避免不必要的等待
await page.goto(url, {'waitUntil': 'networkidle2'})

print("Waiting for target elements...")
try:
    # 直接等待目标元素
    await page.waitForSelector('.tab', {'timeout': 15000})
    elements = await page.querySelectorAll('.tab')
    print(f"找到 {len(elements)} 个class为'tab'的元素")
except Exception as e:
    print(f"查找元素失败: {e}")

await page.screenshot({'path': 'screenshot_TEST.png', 'fullPage': True})

# 获取HTML内容
print("获取HTML内容...")
html = await page.content()
with open("index.html", "w", encoding="utf-8") as file:
   file.write(html)

内容的提问来源于stack exchange,提问作者Samy Rashed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 17:37:24