You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Nodriver进行网页抓取时调用cdp.network.get_response_body卡住的问题排查

Nodriver进行网页抓取时调用cdp.network.get_response_body卡住的问题排查

嘿,我仔细看了你的代码,发现卡住的核心原因是监听的CDP事件时机不对!

你现在用的是RequestWillBeSent事件,这个事件触发的时机是浏览器即将发送请求的时候——这时候服务器还没返回任何响应内容呢!你这时候调用get_response_body,相当于在等一个还没产生的响应,自然会一直卡住,永远等不到结果。

正确的做法:监听ResponseReceived事件

要获取响应体,你需要监听ResponseReceived事件,这个事件是浏览器接收到服务器响应头的时候触发的,这时候响应体已经开始(或者即将)返回,调用get_response_body才能拿到有效数据。

修改后的代码示例

import nodriver
import asyncio
import os

url = "https://www.getfpv.com/graphql?query=%0A%7B%0A%20%20products(%0A%20%20%20%20filter%3A%20%7B%0A%20%20%20%20%20%20%0A%20%20%20%20%20%20sku%3A%20%7B%20in%3A%20%5B%2222817%22%5D%20%7D%2C%0A%20%20%20%20%20%20%0A%20%20%20%20%20%20%0A%20%20%20%20%20%20%7D%0A%20%20%20%20pageSize%3A%208%2C%0A%20%20%20%20sort%3A%20%7Bposition%3A%20ASC%7D%0A%20%20%20%20)%20%7B%0A%20%20%20%20items%20%7B%0A%20%20%20%20%20%20%20%20related_products%20%7B%0A%20%20%20%20%20%20%20%20%0A%20%20%20%20%20%20sku%0A%20%20%20%20%20%20id%0A%20%20%20%20%20%20name%0A%20%20%20%20%20%20small_image%20%7B%0A%20%20%20%20%20%20%20%20label%0A%20%20%20%20%20%20%20%20url%0A%20%20%20%20%20%20%7D%0A%20%20%20%20%20%20thumbnail%20%7B%0A%20%20%20%20%20%20%20%20url%0A%20%20%20%20%20%20%7D%0A%20%20%20%20%20%20url_key%0A%20%20%20%20%20%20url_suffix%0A%20%20%20%20%20%20visibility%0A%20%20%20%20%20%20status%0A%20%20%20%20%20%20price_range%20%7B%0A%20%20%20%20%20%20%20%20minimum_price%20%7B%0A%20%20%20%20%20%20%20%20%20%20regular_price%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20value%0A%20%20%20%20%20%20%20%20%20%20%20%20currency%0A%20%20%20%20%20%20%20%20%20%20%7D%0A%20%20%20%20%20%20%20%20%20%20final_price%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20value%0A%20%20%20%20%20%20%20%20%20%20%20%20currency%0A%20%20%20%20%20%20%20%20%20%20%7D%0A%20%20%20%20%20%20%20%20%7D%0A%20%20%20%20%20%20%7D%0A%20%20%20%20%7D%0A%20%20%7D%0A%20%20%7D%0A%7D"

class Scraper:
    tab: nodriver.Tab

    def __init__(self):
        nodriver.loop().run_until_complete(self.main())

    async def handle_response(self, event):
        print(f"Event: {event}")
        request_id = event.request_id
        print(f"Request ID: {request_id}")
        # 只处理目标URL的响应
        if event.response.url == url:
            print("Trying to get response body...")
            try:
                body = await self.tab.send(nodriver.cdp.network.get_response_body(request_id))
                print(f"Body: {body}")
                # 如果需要解析JSON,可以在这里处理
                # import json
                # parsed_body = json.loads(body['body'])
            except Exception as e:
                print(f"Error getting response body: {e}")

    async def main(self):
        browser = await nodriver.start()
        self.tab = browser.main_tab
        await self.tab.send(nodriver.cdp.network.enable())
        # 替换成监听ResponseReceived事件
        self.tab.add_handler(nodriver.cdp.network.ResponseReceived, self.handle_response)
        await self.tab.get(url)
        await self.tab.wait(t=10)


if __name__ == "__main__":
    Scraper()

额外注意事项

  • 我加了异常捕获,避免因为某些特殊情况(比如响应为空、浏览器未完全接收响应)导致程序崩溃。
  • 如果你的目标是大文件或流式响应,可能需要监听LoadingFinished事件,等响应完全加载完成后再获取body;但对于普通API请求,ResponseReceived就足够用了。
  • 记得通过event.response.url匹配目标请求,避免处理无关的响应。

备注:内容来源于stack exchange,提问作者Fab49er

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 10:55:31