如何使用Python保存网站全量网络流量(含请求响应头与响应体)
问题描述
我想要查找网站加载过程中浏览器下载的对象,目标网站为:https://epco.taleo.net/careersection/alljobs/jobsearch.ftl?lang=en。
我对Web相关技术不太熟悉,想要仅通过网站链接保存请求头、响应头以及实际响应内容。查看网络流量时,能够看到加载末尾有一个jobsearch.ftl?lang=en对象,可查看其响应与头信息。
以下是展示请求与响应头的网络事件日志截图:
以下是实际响应的截图:
以上就是我想要保存的内容,请问该如何实现?
我已尝试如下代码:
import json from selenium import webdriver from selenium.webdriver.common.desired_capabilities import DesiredCapabilities from selenium.webdriver.common.by import By from selenium.common.exceptions import NoSuchElementException, TimeoutException, StaleElementReferenceException from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.chrome.options import Options chromepath = "~/chromedriver/chromedriver" caps = DesiredCapabilities.CHROME caps['goog:loggingPrefs'] = {'performance': 'ALL'} driver = webdriver.Chrome(executable_path=chromepath, desired_capabilities=caps) driver.get('https://epco.taleo.net/careersection/alljobs/jobsearch.ftl?lang=en') def process_browser_log_entry(entry): response = json.loads(entry['message'])['message'] return response browser_log = driver.get_log('performance') events = [process_browser_log_entry(entry) for entry in browser_log] events = [event for event in events if 'Network.response' in event['method']]
但我仅获取到了部分头信息,内容如下:
{'method': 'Network.responseReceivedExtraInfo', 'params': {'blockedCookies': [], 'headers': {'Cache-Control': 'private', 'Connection': 'Keep-Alive', 'Content-Encoding': 'gzip', 'Content-Security-Policy': "frame-ancestors 'self'", 'Content-Type': 'text/html;charset=UTF-8', 'Date': 'Mon, 27 Sep 2021 18:18:10 GMT', 'Keep-Alive': 'timeout=5, max=100', 'P3P': 'CP="CAO PSA OUR"', 'Server': 'Taleo Web Server 8', 'Set-Cookie': 'locale=en; path=/careersection/; secure; HttpOnly', 'Transfer-Encoding': 'chunked', 'Vary': 'Accept-Encoding', 'X-Content-Type-Options': 'nosniff', 'X-UA-Compatible': 'IE=edge', 'X-XSS-Protection': '1'}, 'headersText': 'HTTP/1.1 200 OK\r\nDate: Mon, 27 Sep 2021 18:18:10 GMT\r\nServer: Taleo Web Server 8\r\nCache-Control: private\r\nP3P: CP="CAO PSA OUR"\r\nContent-Encoding: gzip\r\nVary: Accept-Encoding\r\nX-Content-Type-Options: nosniff\r\nSet-Cookie: locale=en; path=/careersection/; secure; HttpOnly\r\nContent-Security-Policy: frame-ancestors \'self\'\r\nX-XSS-Protection: 1\r\nX-UA-Compatible: IE=edge\r\nKeep-Alive: timeout=5, max=100\r\nConnection: Keep-Alive\r\nTransfer-Encoding: chunked\r\nContent-Type: text/html;charset=UTF-8\r\n\r\n', 'requestId': '1E3CDDE80EE37825EF2D9C909FFFAFF3', 'resourceIPAddressSpace': 'Public'}}, {'method': 'Network.responseReceived', 'params': {'frameId': '1624E6F3E724CA508A6D55D556CBE198', 'loaderId': '1E3CDDE80EE37825EF2D9C909FFFAFF3', 'requestId': '1E3CDDE80EE37825EF2D9C909FFFAFF3', 'response': {'connectionId': 26,
这些内容并不包含我在Chrome网页检查器中看到的全部信息。我想要获取完整的请求头、响应头以及实际响应内容,请问上述方法是否正确?有没有无需使用Selenium、仅用requests库的更优方案?
解决方案
原有Selenium方案的问题修正
你之前的方法只过滤了响应接收相关的事件,没有获取请求头数据,也没有调用Chrome DevTools协议的Network.getResponseBody接口获取实际响应内容,所以只拿到了部分信息。可以参考下面的完善代码:
import json from selenium import webdriver from selenium.webdriver.common.desired_capabilities import DesiredCapabilities chromepath = "~/chromedriver/chromedriver" target_url = "https://epco.taleo.net/careersection/alljobs/jobsearch.ftl?lang=en" caps = DesiredCapabilities.CHROME caps['goog:loggingPrefs'] = {'performance': 'ALL'} driver = webdriver.Chrome(executable_path=chromepath, desired_capabilities=caps) driver.get(target_url) # 处理所有性能日志 logs = driver.get_log('performance') events = [json.loads(entry['message'])['message'] for entry in logs] # 找到目标请求的requestId target_request_id = None for event in events: if event['method'] == 'Network.responseReceived': if target_url in event['params']['response']['url']: target_request_id = event['params']['requestId'] # 打印完整响应头 print("=== 响应头 ===") print(json.dumps(event['params']['response']['headers'], indent=2, ensure_ascii=False)) break # 获取请求头 for event in events: if event['method'] == 'Network.requestWillBeSent' and event['params']['requestId'] == target_request_id: print("=== 请求头 ===") print(json.dumps(event['params']['request']['headers'], indent=2, ensure_ascii=False)) break # 获取响应内容 if target_request_id: response_body = driver.execute_cdp_cmd('Network.getResponseBody', {'requestId': target_request_id}) print("=== 响应内容 ===") print(response_body['body']) driver.quit()
纯requests实现方案
这个页面不需要前端动态渲染即可获取返回内容,完全可以直接用requests库实现,操作更简单、性能更好,代码如下:
import requests import json url = "https://epco.taleo.net/careersection/alljobs/jobsearch.ftl?lang=en" # 模拟浏览器UA,避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } resp = requests.get(url, headers=headers) # 保存请求头 with open("request_headers.json", "w", encoding="utf-8") as f: json.dump(dict(resp.request.headers), f, indent=2, ensure_ascii=False) # 保存响应头 with open("response_headers.json", "w", encoding="utf-8") as f: json.dump(dict(resp.headers), f, indent=2, ensure_ascii=False) # 保存响应内容 with open("response_content.html", "w", encoding="utf-8") as f: f.write(resp.text)
运行后三个文件会分别保存你需要的全部内容,不需要启动浏览器,执行效率更高。
内容的提问来源于stack exchange,提问作者anarchy
相关产品推荐
相关产品推荐

