You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+Playwright提取API请求的cURL并打印到终端

问题描述

尝试用Python的Playwright模块拦截HTTP网络请求,获取请求对应的cURL(类似浏览器里复制API调用的cURL格式),但现有代码仅能输出请求对象的字符串表示,无法得到目标cURL文本。

现有尝试代码:

from playwright.sync_api import sync_playwright

def scrape_graphql_response():
    with sync_playwright() as pw:
        browser = pw.chromium.launch(
            headless=False,
            slow_mo=1000,
        )

        context = browser.new_context(
            viewport={"width": 1920, "height": 1080}
        )
        page = context.new_page()

        # Enable request interception
        page.route('**/graphql', lambda route: route.continue_())

        # Navigate to your website
        page.goto('https://example.com')  #sample url for website home page

        # Wait for the GraphQL request to complete
        with page.expect_request('**/graphql') as req:
            # Access the GraphQL response
            print(req.value)

        # Close the browser
        browser.close()


if __name__ == "__main__":
    scrape_graphql_response()
解决方案

要生成cURL命令,需从Playwright的Request对象中提取请求方法、URL、请求头、请求体等核心信息,手动拼接成标准cURL格式即可。

实现代码

编写一个辅助函数处理Request对象到cURL的转换,再在请求拦截逻辑中调用该函数:

from playwright.sync_api import sync_playwright

def request_to_curl(request):
    # 初始化cURL命令基础部分
    curl_cmd = ["curl", "-X", request.method]
    
    # 添加所有请求头
    for name, value in request.headers.items():
        curl_cmd.extend(["-H", f'"{name}: {value}"'])
    
    # 处理POST/PUT等带请求体的请求
    if request.post_data:
        content_type = request.headers.get("Content-Type", "")
        # 针对JSON格式请求体做特殊处理
        if "application/json" in content_type:
            curl_cmd.extend(["-d", f'"{request.post_data}"'])
        else:
            curl_cmd.extend(["-d", request.post_data])
    
    # 添加目标URL
    curl_cmd.append(f'"{request.url}"')
    
    # 拼接成完整的cURL命令字符串
    return " ".join(curl_cmd)

def scrape_graphql_response():
    with sync_playwright() as pw:
        browser = pw.chromium.launch(
            headless=False,
            slow_mo=1000,
        )

        context = browser.new_context(
            viewport={"width": 1920, "height": 1080}
        )
        page = context.new_page()

        # 拦截GraphQL请求并生成cURL
        def handle_graphql_request(route):
            curl_command = request_to_curl(route.request)
            print("生成的cURL命令:")
            print(curl_command)
            route.continue_()

        page.route('**/graphql', handle_graphql_request)

        # 导航到目标网站
        page.goto('https://example.com')

        # 等待请求完成,确保捕获到目标请求
        page.wait_for_request('**/graphql')

        browser.close()


if __name__ == "__main__":
    scrape_graphql_response()

关键说明

  • 借助route.request获取拦截到的请求对象,提取method、url、headers、post_data等信息构建cURL。
  • 对包含特殊字符的请求头、请求体和URL用双引号包裹,避免命令执行出错。
  • 可根据实际需求扩展处理逻辑,比如添加--insecure忽略SSL验证、处理多部分表单数据等。

内容的提问来源于stack exchange,提问作者Ashutosh Yadav

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 22:05:18