如何在Scrapy Playwright中避免选择器未找到引发的页面超时崩溃?
解决Scrapy Playwright中GDPR按钮不存在时的超时问题
原代码中使用playwright_page_methods里的wait_for_selector会在GDPR按钮不存在时触发超时错误,因为该方法会默认等待元素出现直到超时。要实现「元素存在则点击,不存在则跳过」的逻辑,需要用自定义页面回调函数替代固定的PageMethod序列。
修改后的完整代码
from typing import Iterable import scrapy from playwright.async_api import Page GDPR_BUTTON_SELECTOR = "iframe[id^='sp_message_iframe'] >> internal:control=enter-frame >> .sp_choice_type_11" class GuardianSpider(scrapy.Spider): name = "guardian" allowed_domains = ["www.theguardian.com"] start_urls = ["https://www.theguardian.com"] async def handle_gdpr(self, page: Page) -> None: # 检查GDPR按钮是否存在 if await page.locator(GDPR_BUTTON_SELECTOR).count() > 0: await page.locator(GDPR_BUTTON_SELECTOR).dispatch_event("click") # 可选:等待页面状态稳定后再继续 await page.wait_for_timeout(1000) def start_requests(self) -> Iterable[scrapy.Request]: url = "https://www.theguardian.com" yield scrapy.Request( url, meta=dict( playwright=True, # 指定自定义页面处理回调 playwright_page_callback=self.handle_gdpr, ), ) def parse(self, response): # 此处添加你的响应处理逻辑 pass
关键说明
使用
playwright_page_callback替代playwright_page_methods
该回调函数允许编写自定义异步逻辑,直接操作Playwright的Page对象,完全复刻原生Playwright的条件判断逻辑。异步函数要求
回调函数必须用async def定义,因为Scrapy Playwright基于异步Playwright运行,所有页面操作都需要用await关键字。元素存在性判断
通过await page.locator(selector).count()获取元素数量,大于0则说明元素存在,执行点击操作;否则直接跳过。可选优化
点击按钮后添加wait_for_timeout,确保页面有足够时间处理GDPR确认后的状态变化(如弹窗关闭、页面刷新等)。
替代实现(基于异常捕获)
如果偏好使用wait_for_selector,可以通过捕获超时异常实现相同逻辑:
async def handle_gdpr(self, page: Page) -> None: from playwright._impl._errors import TimeoutError try: # 设置超时为0,不等待直接检查元素是否存在 await page.wait_for_selector(GDPR_BUTTON_SELECTOR, timeout=0) await page.locator(GDPR_BUTTON_SELECTOR).dispatch_event("click") except TimeoutError: # 元素不存在,跳过操作 pass
内容的提问来源于stack exchange,提问作者muw
相关产品推荐
相关产品推荐

