You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scrapy-Playwright中监听page.on并捕获POST请求的JSON响应

如何在Scrapy-Playwright中捕获指定POST请求的响应并传递到parse函数

你可以通过两种方式实现目标:注册全局响应监听器,或者直接等待目标响应。以下是具体实现方案:

方法一:通过页面初始化回调注册响应监听器

利用Scrapy-Playwright提供的playwright_page_init_callback,在Page实例创建时注册response事件监听,捕获目标请求的响应并存储到页面属性中,后续在parse函数中读取。

import asyncio
import scrapy

class YourSpider(scrapy.Spider):
    name = "turo_spider"
    start_urls = ["https://turo.com/gb/en/search?country=US&..."]  # 替换为你的起始URL

    def start_requests(self):
        def init_page(page):
            # 定义响应处理函数
            async def capture_target_response(response):
                # 匹配目标POST请求
                if (response.request.method == "POST" 
                    and response.url == "https://turo.com/api/bulk-quotes/v2"):
                    # 将响应JSON存储到page的自定义属性
                    page.target_json = await response.json()
            
            # 注册响应监听事件
            page.on("response", capture_target_response)

        for url in self.start_urls:
            yield scrapy.Request(
                url,
                meta={
                    "playwright": True,
                    "playwright_include_page": True,
                    "playwright_page_init_callback": init_page,  # 绑定页面初始化回调
                },
            )

    async def parse(self, response):
        page = response.meta["playwright_page"]
        # 等待目标响应被捕获(根据实际情况调整等待逻辑)
        while not hasattr(page, "target_json"):
            await asyncio.sleep(0.1)
        
        # 获取捕获到的JSON数据
        target_data = page.target_json
        self.logger.info("成功捕获目标POST响应")
        # 这里添加你的后续处理逻辑
        # ...

方法二:直接等待目标响应(更简洁)

在parse函数中使用Playwright原生的page.wait_for_response方法,直接等待符合条件的响应出现,无需全局监听。

import scrapy

class YourSpider(scrapy.Spider):
    name = "turo_spider"
    start_urls = ["https://turo.com/gb/en/search?country=US&..."]  # 替换为你的起始URL

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(
                url,
                meta={
                    "playwright": True,
                    "playwright_include_page": True,
                },
            )

    async def parse(self, response):
        page = response.meta["playwright_page"]
        # 等待目标POST请求的响应,lambda函数用于匹配请求
        target_response = await page.wait_for_response(
            lambda resp: resp.request.method == "POST" and resp.url == "https://turo.com/api/bulk-quotes/v2"
        )
        # 解析响应为JSON格式
        target_data = await target_response.json()
        self.logger.info("成功捕获目标POST响应")
        # 后续处理逻辑
        # ...

注意事项

  • parse函数必须定义为async函数,因为需要调用Playwright的异步方法。
  • 如果目标POST请求是由页面交互(如点击按钮、滚动)触发的,需要在等待响应前先执行对应的交互操作(比如await page.click(selector))。
  • 确保你的Scrapy-Playwright版本支持上述特性,建议使用最新稳定版。

内容的提问来源于stack exchange,提问作者Onyilimba

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 09:25:03