You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Selenium搭配mitmproxy时拦截指定URL资源?

问题:Selenium搭配mitmproxy无法拦截指定URL资源

已配置mitmproxy对接上游代理,期望流量流程为:Python主程序→mitmproxy(拦截目标资源,不传递到上游)→上游代理→网页,返回流程反向。但目前部分URL(如gvt1.com、statically.io、jsdelivr.net、gstatic.com、easylist.to等)未按预期被拦截。

已使用的mitmproxy启动命令:

mitmproxy --mode upstream:https://HOSTNAME:PORT --upstream-auth USER:PASSWORD

原block.py代码:

# block.py
from mitmproxy import http

# we can block popular 3rd party resources like tracking and advertisements.
BLOCK_RESOURCE_NAMES = [
  'adzerk',
  'analytics',
  'cdn.api.twitter',
  'doubleclick',
  'exelator',
  'facebook',
  'fontawesome',
  'google',
  'google-analytics',
  'googletagmanager',
  # or something abstract like images
  'images',
  'gvt1'
]
# or block based on resource extension
BLOCK_RESOURCE_EXTENSIONS = [
    '.gif',
    '.jpg',
    '.jpeg',
    '.png',
    '.webp',
    '.com'
]

# this will handle all requests going through proxy:
def request(flow: http.HTTPFlow) -> None:
    url = flow.request.pretty_url
    has_blocked_extension = any(url.endswith(ext) for ext in BLOCK_RESOURCE_EXTENSIONS)
    contains_blocked_key = any(block in url for block in BLOCK_RESOURCE_NAMES)
    if has_blocked_extension or contains_blocked_key:
        print(f"Blocked {url}")
        flow.response = http.Response.make(
            404,  # status code
            b"Blocked",  # content
            {"Content-Type": "text/html"}  # headers
        )

原Python主文件代码:

# I want to have following flow

#python_main_file -> mitmproxy(if we can block all the resources here, so that
# it won't pass to other proxy) -> otherproxy -> webpage and flow back the same pattern

from selenium import webdriver
import subprocess 

command = f"mitmdump -s block.py --mode upstream:https://HOSTNAME:PORT --upstream-auth USER:PASSWORD"
subprocess.Popen(command, shell=True)

PROXY = "localhost:8080"  # IP:PORT or HOST:PORT of our mitmproxy

chrome_options = webdriver.ChromeOptions()
# this command enabled proxy for our Selenium browser:
chrome_options.add_argument('--proxy-server=%s' % PROXY)

chrome = webdriver.Chrome(options=chrome_options)
# test it by going to a page with blocked resources:
# chrome.get("https://web-scraping.dev/product/1")
chrome.get("https://www.whatsmyip.org/")
chrome.quit()

解决方案

1. 修正拦截规则的核心问题

原代码存在两个关键缺陷:

  • 拦截关键词未覆盖目标域名:BLOCK_RESOURCE_NAMES仅包含gvt1,未添加statically、jsdelivr、gstatic、easylist等需要拦截的域名关键词
  • 扩展名拦截逻辑错误:.com作为扩展名不合理,URL不会以.com结尾(后续还有路径/参数),且该规则会误拦截大量正常请求

建议直接通过请求的Host字段进行域名匹配,比URL包含匹配更精准,避免路径/参数干扰。

2. 修正后的block.py代码

# block.py
from mitmproxy import http

# 要拦截的目标域名关键词(包含子域名)
BLOCK_DOMAIN_KEYWORDS = [
    'adzerk',
    'analytics',
    'cdn.api.twitter',
    'doubleclick',
    'exelator',
    'facebook',
    'fontawesome',
    'google',
    'google-analytics',
    'googletagmanager',
    'gvt1',
    'statically',
    'jsdelivr',
    'gstatic',
    'easylist'
]

# 要拦截的资源扩展名
BLOCK_RESOURCE_EXTENSIONS = [
    '.gif',
    '.jpg',
    '.jpeg',
    '.png',
    '.webp'
]

def request(flow: http.HTTPFlow) -> None:
    # 获取请求的Host(纯域名部分)
    host = flow.request.host
    url = flow.request.pretty_url
    
    # 检查域名是否包含拦截关键词
    domain_match = any(keyword in host for keyword in BLOCK_DOMAIN_KEYWORDS)
    # 检查URL是否以指定扩展名结尾
    ext_match = any(url.endswith(ext) for ext in BLOCK_RESOURCE_EXTENSIONS)
    
    if domain_match or ext_match:
        print(f"Blocked {url}")
        flow.response = http.Response.make(
            404,
            b"Blocked",
            {"Content-Type": "text/html"}
        )

3. 补充配置注意事项

  • 信任mitmproxy证书:Chrome浏览器需要信任mitmproxy的CA证书,否则HTTPS请求会被拦截失败。启动mitmproxy后访问http://mitm.it下载对应系统的证书,在Chrome设置中导入并设为信任。
  • 确保mitmproxy先就绪:主程序中启动mitmproxy后,添加time.sleep(2)短暂延迟,避免Selenium连接时mitmproxy尚未完全启动。
  • 验证上游代理连通性:单独用mitmproxy测试上游代理的地址和认证信息是否正确,确保流量能正常转发。

内容的提问来源于stack exchange,提问作者Bad Coder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 14:03:21