Python中使用Browsermobproxy+Selenium+FireFox无法获取HAR响应体
解决Browsermobproxy+Selenium+Firefox抓包时HAR无响应体的问题
我之前踩过一模一样的坑!结合你的代码,咱们从几个关键地方排查修复:
1. 必须配置HTTPS证书信任
Browsermobproxy抓HTTPS请求时会生成自签名证书,Firefox默认不信任,这会直接导致响应内容无法被捕获。你需要把代理证书导入到Firefox Profile中:
# 在创建profile后添加这行 profile.add_extension(proxy.certificate)
2. 强化抓包参数(全局+HAR实例)
虽然你已经在创建proxy时设置了抓包参数,但有时候全局设置可能不生效,建议在new_har时重复指定一次,确保覆盖所有配置:
proxy.new_har('xxx', options={ 'captureHeaders': True, 'captureContent': True, 'captureBinaryContent': True })
3. 禁用Firefox缓存避免干扰
缓存会导致部分请求直接从本地读取,完全不经过代理,自然不会被抓包工具捕获。添加这些偏好设置禁用缓存:
# 在设置proxy之前添加这些配置 profile.set_preference("network.http.use-cache", False) profile.set_preference("browser.cache.disk.enable", False) profile.set_preference("browser.cache.memory.enable", False) profile.set_preference("browser.cache.offline.enable", False)
4. 确保请求完全完成后再获取HAR
你的代码里driver.get后直接获取HAR,大概率页面请求还没全部完成,导致部分响应没被记录。建议添加等待逻辑:
driver.get('XXX') # 要么用固定等待(简单但不够严谨) time.sleep(5) # 要么用显式等待某个页面元素加载(更可靠) from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.TAG_NAME, "body")) )
修改后的完整代码示例
from selenium import webdriver from browsermobproxy import Server from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import json import time server = Server(r'D:\browsermob-proxy-2.1.4\bin\browsermob-proxy.bat') server.start() proxy = server.create_proxy({ 'captureHeaders': True, 'captureContent': True, 'captureBinaryContent': True }) # 配置Firefox Profile profile = webdriver.FirefoxProfile() # 禁用缓存 profile.set_preference("network.http.use-cache", False) profile.set_preference("browser.cache.disk.enable", False) profile.set_preference("browser.cache.memory.enable", False) profile.set_preference("browser.cache.offline.enable", False) # 导入代理证书 profile.add_extension(proxy.certificate) # 设置代理 profile.set_proxy(proxy.selenium_proxy()) # 注意:新版本Firefox建议用options参数替代firefox_profile from selenium.webdriver.firefox.options import Options options = Options() options.profile = profile driver = webdriver.Firefox(options=options) # 创建HAR时重复指定抓包参数 proxy.new_har('xxx', options={ 'captureHeaders': True, 'captureContent': True, 'captureBinaryContent': True }) driver.get('XXX') # 等待页面加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.TAG_NAME, "body")) ) # 获取并保存HAR文件 result = proxy.har with open('output.har', 'w', encoding='utf-8') as f: json.dump(result, f, indent=2) # 清理资源 driver.quit() server.stop()
内容的提问来源于stack exchange,提问作者chbr
相关产品推荐
相关产品推荐

