如何用Python+Playwright提取API请求的cURL并打印到终端
问题描述
尝试用Python的Playwright模块拦截HTTP网络请求,获取请求对应的cURL(类似浏览器里复制API调用的cURL格式),但现有代码仅能输出请求对象的字符串表示,无法得到目标cURL文本。
现有尝试代码:
from playwright.sync_api import sync_playwright def scrape_graphql_response(): with sync_playwright() as pw: browser = pw.chromium.launch( headless=False, slow_mo=1000, ) context = browser.new_context( viewport={"width": 1920, "height": 1080} ) page = context.new_page() # Enable request interception page.route('**/graphql', lambda route: route.continue_()) # Navigate to your website page.goto('https://example.com') #sample url for website home page # Wait for the GraphQL request to complete with page.expect_request('**/graphql') as req: # Access the GraphQL response print(req.value) # Close the browser browser.close() if __name__ == "__main__": scrape_graphql_response()
解决方案
要生成cURL命令,需从Playwright的Request对象中提取请求方法、URL、请求头、请求体等核心信息,手动拼接成标准cURL格式即可。
实现代码
编写一个辅助函数处理Request对象到cURL的转换,再在请求拦截逻辑中调用该函数:
from playwright.sync_api import sync_playwright def request_to_curl(request): # 初始化cURL命令基础部分 curl_cmd = ["curl", "-X", request.method] # 添加所有请求头 for name, value in request.headers.items(): curl_cmd.extend(["-H", f'"{name}: {value}"']) # 处理POST/PUT等带请求体的请求 if request.post_data: content_type = request.headers.get("Content-Type", "") # 针对JSON格式请求体做特殊处理 if "application/json" in content_type: curl_cmd.extend(["-d", f'"{request.post_data}"']) else: curl_cmd.extend(["-d", request.post_data]) # 添加目标URL curl_cmd.append(f'"{request.url}"') # 拼接成完整的cURL命令字符串 return " ".join(curl_cmd) def scrape_graphql_response(): with sync_playwright() as pw: browser = pw.chromium.launch( headless=False, slow_mo=1000, ) context = browser.new_context( viewport={"width": 1920, "height": 1080} ) page = context.new_page() # 拦截GraphQL请求并生成cURL def handle_graphql_request(route): curl_command = request_to_curl(route.request) print("生成的cURL命令:") print(curl_command) route.continue_() page.route('**/graphql', handle_graphql_request) # 导航到目标网站 page.goto('https://example.com') # 等待请求完成,确保捕获到目标请求 page.wait_for_request('**/graphql') browser.close() if __name__ == "__main__": scrape_graphql_response()
关键说明
- 借助
route.request获取拦截到的请求对象,提取method、url、headers、post_data等信息构建cURL。 - 对包含特殊字符的请求头、请求体和URL用双引号包裹,避免命令执行出错。
- 可根据实际需求扩展处理逻辑,比如添加
--insecure忽略SSL验证、处理多部分表单数据等。
内容的提问来源于stack exchange,提问作者Ashutosh Yadav
相关产品推荐
相关产品推荐

