You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python保存网站全量网络流量(含请求响应头与响应体)

问题描述

我想要查找网站加载过程中浏览器下载的对象,目标网站为:https://epco.taleo.net/careersection/alljobs/jobsearch.ftl?lang=en。
我对Web相关技术不太熟悉,想要仅通过网站链接保存请求头、响应头以及实际响应内容。查看网络流量时,能够看到加载末尾有一个jobsearch.ftl?lang=en对象,可查看其响应与头信息。
以下是展示请求与响应头的网络事件日志截图:
network event log
以下是实际响应的截图:
response

以上就是我想要保存的内容,请问该如何实现?
我已尝试如下代码:

import json
from selenium import webdriver
from selenium.webdriver.common.desired_capabilities import DesiredCapabilities
from selenium.webdriver.common.by import By
from selenium.common.exceptions import NoSuchElementException, TimeoutException, StaleElementReferenceException
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.options import Options

chromepath = "~/chromedriver/chromedriver"

caps = DesiredCapabilities.CHROME
caps['goog:loggingPrefs'] = {'performance': 'ALL'}
driver = webdriver.Chrome(executable_path=chromepath, desired_capabilities=caps)
driver.get('https://epco.taleo.net/careersection/alljobs/jobsearch.ftl?lang=en')

def process_browser_log_entry(entry):
    response = json.loads(entry['message'])['message']
    return response

browser_log = driver.get_log('performance') 
events = [process_browser_log_entry(entry) for entry in browser_log]
events = [event for event in events if 'Network.response' in event['method']]

但我仅获取到了部分头信息,内容如下:

{'method': 'Network.responseReceivedExtraInfo',
  'params': {'blockedCookies': [],
   'headers': {'Cache-Control': 'private',
    'Connection': 'Keep-Alive',
    'Content-Encoding': 'gzip',
    'Content-Security-Policy': "frame-ancestors 'self'",
    'Content-Type': 'text/html;charset=UTF-8',
    'Date': 'Mon, 27 Sep 2021 18:18:10 GMT',
    'Keep-Alive': 'timeout=5, max=100',
    'P3P': 'CP="CAO PSA OUR"',
    'Server': 'Taleo Web Server 8',
    'Set-Cookie': 'locale=en; path=/careersection/; secure; HttpOnly',
    'Transfer-Encoding': 'chunked',
    'Vary': 'Accept-Encoding',
    'X-Content-Type-Options': 'nosniff',
    'X-UA-Compatible': 'IE=edge',
    'X-XSS-Protection': '1'},
   'headersText': 'HTTP/1.1 200 OK\r\nDate: Mon, 27 Sep 2021 18:18:10 GMT\r\nServer: Taleo Web Server 8\r\nCache-Control: private\r\nP3P: CP="CAO PSA OUR"\r\nContent-Encoding: gzip\r\nVary: Accept-Encoding\r\nX-Content-Type-Options: nosniff\r\nSet-Cookie: locale=en; path=/careersection/; secure; HttpOnly\r\nContent-Security-Policy: frame-ancestors \'self\'\r\nX-XSS-Protection: 1\r\nX-UA-Compatible: IE=edge\r\nKeep-Alive: timeout=5, max=100\r\nConnection: Keep-Alive\r\nTransfer-Encoding: chunked\r\nContent-Type: text/html;charset=UTF-8\r\n\r\n',
   'requestId': '1E3CDDE80EE37825EF2D9C909FFFAFF3',
   'resourceIPAddressSpace': 'Public'}},
 {'method': 'Network.responseReceived',
  'params': {'frameId': '1624E6F3E724CA508A6D55D556CBE198',
   'loaderId': '1E3CDDE80EE37825EF2D9C909FFFAFF3',
   'requestId': '1E3CDDE80EE37825EF2D9C909FFFAFF3',
   'response': {'connectionId': 26,

这些内容并不包含我在Chrome网页检查器中看到的全部信息。我想要获取完整的请求头、响应头以及实际响应内容,请问上述方法是否正确?有没有无需使用Selenium、仅用requests库的更优方案?

解决方案

原有Selenium方案的问题修正

你之前的方法只过滤了响应接收相关的事件,没有获取请求头数据,也没有调用Chrome DevTools协议的Network.getResponseBody接口获取实际响应内容,所以只拿到了部分信息。可以参考下面的完善代码:

import json
from selenium import webdriver
from selenium.webdriver.common.desired_capabilities import DesiredCapabilities

chromepath = "~/chromedriver/chromedriver"
target_url = "https://epco.taleo.net/careersection/alljobs/jobsearch.ftl?lang=en"

caps = DesiredCapabilities.CHROME
caps['goog:loggingPrefs'] = {'performance': 'ALL'}
driver = webdriver.Chrome(executable_path=chromepath, desired_capabilities=caps)
driver.get(target_url)

# 处理所有性能日志
logs = driver.get_log('performance')
events = [json.loads(entry['message'])['message'] for entry in logs]

# 找到目标请求的requestId
target_request_id = None
for event in events:
    if event['method'] == 'Network.responseReceived':
        if target_url in event['params']['response']['url']:
            target_request_id = event['params']['requestId']
            # 打印完整响应头
            print("=== 响应头 ===")
            print(json.dumps(event['params']['response']['headers'], indent=2, ensure_ascii=False))
            break

# 获取请求头
for event in events:
    if event['method'] == 'Network.requestWillBeSent' and event['params']['requestId'] == target_request_id:
        print("=== 请求头 ===")
        print(json.dumps(event['params']['request']['headers'], indent=2, ensure_ascii=False))
        break

# 获取响应内容
if target_request_id:
    response_body = driver.execute_cdp_cmd('Network.getResponseBody', {'requestId': target_request_id})
    print("=== 响应内容 ===")
    print(response_body['body'])

driver.quit()

纯requests实现方案

这个页面不需要前端动态渲染即可获取返回内容,完全可以直接用requests库实现,操作更简单、性能更好,代码如下:

import requests
import json

url = "https://epco.taleo.net/careersection/alljobs/jobsearch.ftl?lang=en"
# 模拟浏览器UA,避免被反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

resp = requests.get(url, headers=headers)

# 保存请求头
with open("request_headers.json", "w", encoding="utf-8") as f:
    json.dump(dict(resp.request.headers), f, indent=2, ensure_ascii=False)

# 保存响应头
with open("response_headers.json", "w", encoding="utf-8") as f:
    json.dump(dict(resp.headers), f, indent=2, ensure_ascii=False)

# 保存响应内容
with open("response_content.html", "w", encoding="utf-8") as f:
    f.write(resp.text)

运行后三个文件会分别保存你需要的全部内容,不需要启动浏览器,执行效率更高。

内容的提问来源于stack exchange,提问作者anarchy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 06:54:03