如何在Selenium搭配mitmproxy时拦截指定URL资源?
问题:Selenium搭配mitmproxy无法拦截指定URL资源
已配置mitmproxy对接上游代理,期望流量流程为:Python主程序→mitmproxy(拦截目标资源,不传递到上游)→上游代理→网页,返回流程反向。但目前部分URL(如gvt1.com、statically.io、jsdelivr.net、gstatic.com、easylist.to等)未按预期被拦截。
已使用的mitmproxy启动命令:
mitmproxy --mode upstream:https://HOSTNAME:PORT --upstream-auth USER:PASSWORD
原block.py代码:
# block.py from mitmproxy import http # we can block popular 3rd party resources like tracking and advertisements. BLOCK_RESOURCE_NAMES = [ 'adzerk', 'analytics', 'cdn.api.twitter', 'doubleclick', 'exelator', 'facebook', 'fontawesome', 'google', 'google-analytics', 'googletagmanager', # or something abstract like images 'images', 'gvt1' ] # or block based on resource extension BLOCK_RESOURCE_EXTENSIONS = [ '.gif', '.jpg', '.jpeg', '.png', '.webp', '.com' ] # this will handle all requests going through proxy: def request(flow: http.HTTPFlow) -> None: url = flow.request.pretty_url has_blocked_extension = any(url.endswith(ext) for ext in BLOCK_RESOURCE_EXTENSIONS) contains_blocked_key = any(block in url for block in BLOCK_RESOURCE_NAMES) if has_blocked_extension or contains_blocked_key: print(f"Blocked {url}") flow.response = http.Response.make( 404, # status code b"Blocked", # content {"Content-Type": "text/html"} # headers )
原Python主文件代码:
# I want to have following flow #python_main_file -> mitmproxy(if we can block all the resources here, so that # it won't pass to other proxy) -> otherproxy -> webpage and flow back the same pattern from selenium import webdriver import subprocess command = f"mitmdump -s block.py --mode upstream:https://HOSTNAME:PORT --upstream-auth USER:PASSWORD" subprocess.Popen(command, shell=True) PROXY = "localhost:8080" # IP:PORT or HOST:PORT of our mitmproxy chrome_options = webdriver.ChromeOptions() # this command enabled proxy for our Selenium browser: chrome_options.add_argument('--proxy-server=%s' % PROXY) chrome = webdriver.Chrome(options=chrome_options) # test it by going to a page with blocked resources: # chrome.get("https://web-scraping.dev/product/1") chrome.get("https://www.whatsmyip.org/") chrome.quit()
解决方案
1. 修正拦截规则的核心问题
原代码存在两个关键缺陷:
- 拦截关键词未覆盖目标域名:
BLOCK_RESOURCE_NAMES仅包含gvt1,未添加statically、jsdelivr、gstatic、easylist等需要拦截的域名关键词 - 扩展名拦截逻辑错误:
.com作为扩展名不合理,URL不会以.com结尾(后续还有路径/参数),且该规则会误拦截大量正常请求
建议直接通过请求的Host字段进行域名匹配,比URL包含匹配更精准,避免路径/参数干扰。
2. 修正后的block.py代码
# block.py from mitmproxy import http # 要拦截的目标域名关键词(包含子域名) BLOCK_DOMAIN_KEYWORDS = [ 'adzerk', 'analytics', 'cdn.api.twitter', 'doubleclick', 'exelator', 'facebook', 'fontawesome', 'google', 'google-analytics', 'googletagmanager', 'gvt1', 'statically', 'jsdelivr', 'gstatic', 'easylist' ] # 要拦截的资源扩展名 BLOCK_RESOURCE_EXTENSIONS = [ '.gif', '.jpg', '.jpeg', '.png', '.webp' ] def request(flow: http.HTTPFlow) -> None: # 获取请求的Host(纯域名部分) host = flow.request.host url = flow.request.pretty_url # 检查域名是否包含拦截关键词 domain_match = any(keyword in host for keyword in BLOCK_DOMAIN_KEYWORDS) # 检查URL是否以指定扩展名结尾 ext_match = any(url.endswith(ext) for ext in BLOCK_RESOURCE_EXTENSIONS) if domain_match or ext_match: print(f"Blocked {url}") flow.response = http.Response.make( 404, b"Blocked", {"Content-Type": "text/html"} )
3. 补充配置注意事项
- 信任mitmproxy证书:Chrome浏览器需要信任mitmproxy的CA证书,否则HTTPS请求会被拦截失败。启动mitmproxy后访问
http://mitm.it下载对应系统的证书,在Chrome设置中导入并设为信任。 - 确保mitmproxy先就绪:主程序中启动mitmproxy后,添加
time.sleep(2)短暂延迟,避免Selenium连接时mitmproxy尚未完全启动。 - 验证上游代理连通性:单独用mitmproxy测试上游代理的地址和认证信息是否正确,确保流量能正常转发。
内容的提问来源于stack exchange,提问作者Bad Coder
相关产品推荐
相关产品推荐

