Pyppeteer加载页面后HTML空白,元素选择超时问题排查
我正在使用Pyppeteer(Puppeteer的非官方Python移植版本)进行网页抓取,尝试选择元素时遇到问题。例如,我执行elements = await page.querySelectorAll('.tab')等待class为"tab"的元素,但始终返回空且触发超时错误。实际上,等待该网站任意class的元素都会超时。我尝试导出页面HTML排查,发现不仅目标元素缺失,打开HTML文件时页面完全空白。以下是我的代码片段,请问代码执行逻辑是否存在缺陷?
import asyncio from pyppeteer import launch import requests import sys import json print("Test") print("Starting script...") def print_cookies_as_json(cookies): cookies_json = json.dumps(cookies, indent=4) print("Cookies in JSON format:") print(cookies_json) async def main(): try: print("Launching browser...") browser = await launch(headless=False) page = await browser.newPage() except Exception as e: print(f"Error launching browser or creating new page: {e}") return try: print("Reading cookies from file...") # Load cookies from JSON file try: with open('/path/to/usrCookies.json', 'r') as f: cookies = json.load(f) if not cookies: raise ValueError("No valid cookies found in the file.") await page.setCookie(*cookies) except (FileNotFoundError, ValueError) as e: print(f"Error reading cookies: {e}") return # Print cookies in JSON format print_cookies_as_json(cookies) print("Navigating to URL...") url = 'https://example.com' await page.goto(url, {'waitUntil': 'networkidle0'}) response = requests.get(url, cookies={c['name']: c['value'] for c in cookies}) json_response = response.json() print(json_response) print("Processing JSON data...") await page.screenshot({'path': 'screenshot_TEST.png', 'fullPage': True}) print("Waiting for page to load...") await page.waitForSelector('body', {'timeout': 10000}) # wait for the body to load await asyncio.sleep(1) # Get the HTML content of the page print("Getting HTML content...") html = await page.content() # Write the HTML content to a file with open("index.html", "w", encoding="utf-8") as file: file.write(html) except Exception as e: print("An error occurred:", e) await page.screenshot({'path': 'screenshot_ERROR.png', 'fullPage': True}) print("Screenshot saved as screenshot.png") finally: # Close the browser try: print("Closing browser...") await browser.close() except Exception as e: print(f"Error closing browser: {e}") asyncio.get_event_loop().run_until_complete(main())
多余的requests请求破坏会话:你在
page.goto之后调用了requests.get,这会创建一个独立于Pyppeteer浏览器的新HTTP会话,不仅完全没用,还可能触发网站的反爬机制,导致浏览器中的会话失效,页面被清空。直接删掉这部分代码即可,Pyppeteer已经通过page.setCookie配置了会话Cookie。Cookie导入可能不完整:Pyppeteer的
page.setCookie要求每个Cookie必须包含domain字段(比如".example.com"),如果你的usrCookies.json里没有这个字段,Cookie设置会失效,导致访问网站时没有正确的身份验证,返回空白页面。检查Cookie文件,确保每个Cookie对象都有domain、name、value这些必填项。等待逻辑顺序混乱:
page.goto已经设置了waitUntil: 'networkidle0',之后再等待body元素完全多余。而且如果中间插入了干扰操作(比如那个requests请求),页面状态已经异常,此时等待body毫无意义。应该直接在page.goto之后等待目标元素,比如await page.waitForSelector('.tab', {'timeout': 15000})。固定sleep不可靠:用
asyncio.sleep(1)等待页面加载是碰运气的做法,换成等待目标元素的方式,能确保元素加载完成后再执行后续操作。
修改后的核心代码示例:
print("Navigating to URL...") url = 'https://example.com' # 确保Cookie包含domain字段,否则setCookie不生效 await page.setCookie(*cookies) # networkidle2比networkidle0更适合大多数场景,避免不必要的等待 await page.goto(url, {'waitUntil': 'networkidle2'}) print("Waiting for target elements...") try: # 直接等待目标元素 await page.waitForSelector('.tab', {'timeout': 15000}) elements = await page.querySelectorAll('.tab') print(f"找到 {len(elements)} 个class为'tab'的元素") except Exception as e: print(f"查找元素失败: {e}") await page.screenshot({'path': 'screenshot_TEST.png', 'fullPage': True}) # 获取HTML内容 print("获取HTML内容...") html = await page.content() with open("index.html", "w", encoding="utf-8") as file: file.write(html)
内容的提问来源于stack exchange,提问作者Samy Rashed

