如何修改mitmproxy脚本在HTTP请求阶段拦截多媒体等资源?
问题场景与需求
我希望在浏览任意域名时拦截所有图片、视频、音频等资源,当前mitmdump运行在Linux机器,Windows端Firefox已配置使用该Linux机器的8080端口作为代理,代理环境正常。
最初编写了如下request阶段的mitmproxy脚本,但执行mitmdump -s reject.py后访问cnn.com仍能看到图片,脚本未生效:
from mitmproxy import http def request(flow): if flow.request.headers.get("content-type") == "image/*": flow.response = http.Response.make(404, b"Rejected")
之后改用response阶段的脚本实现了拦截需求:
from mitmproxy import http def response(flow): if flow.response.status_code not in [101, 201, 204, 302, 301, 304]: # 避免content.split出现None类型 my_content = flow.response.headers.get("content-type") if my_content: iana_header = my_content.split('/', 2) if iana_header[0] in ("audio", "example", "font", "image", "message", "model", "video"): flow.response = http.Response.make(404, b"Rejected") elif iana_header[0] == 'application': print(iana_header[1]) # [wordperfect5.1] - 唤起回忆 if iana_header[1] not in ["dns", "gzip", "http", "javascript", "json", "pdf", "rtf", "xhtml+xml", "xml", "xml-dtd", "xml-external-parsed-entity", "xml-patch+xml", "zip"]: flow.response = http.Response.make(404, b"Rejected")
现在需要将拦截逻辑改为在HTTP请求阶段执行,且无需逐个检查文件扩展名,请问该如何修改脚本?
解决方案
原request脚本无效原因
原脚本的核心错误在于:请求头的Content-Type是客户端告知服务器自身发送内容的类型,而图片、视频这类资源的请求多为GET请求,请求头里根本不存在Content-Type字段;即便存在,也不会是image/*这种用于标识响应内容的格式。
修改后的request阶段脚本
结合请求阶段的特征(Accept头、URL路径),同时对齐你原response脚本的拦截范围,编写如下脚本:
from mitmproxy import http import re # 定义需要拦截的顶级媒体类型 BLOCKED_TOP_LEVEL_TYPES = {"audio", "example", "font", "image", "message", "model", "video"} # 定义允许通过的application子类型 ALLOWED_APP_SUBTYPES = {"dns", "gzip", "http", "javascript", "json", "pdf", "rtf", "xhtml+xml", "xml", "xml-dtd", "xml-external-parsed-entity", "xml-patch+xml", "zip"} def request(flow): # 优先通过Accept头判断请求目标类型 accept_header = flow.request.headers.get("Accept", "") if accept_header: # 拆分Accept头中多个媒体类型,忽略权重参数 accept_types = [t.strip().split(';')[0] for t in accept_header.split(',')] for accept_type in accept_types: if '/' not in accept_type: continue top_level, sub_type = accept_type.split('/', 1) # 拦截指定顶级类型的请求 if top_level in BLOCKED_TOP_LEVEL_TYPES: flow.response = http.Response.make(404, b"Rejected") return # 处理application类型,仅允许指定子类型 if top_level == "application": if sub_type == "*" or sub_type not in ALLOWED_APP_SUBTYPES: flow.response = http.Response.make(404, b"Rejected") return # 补充拦截无Accept头但URL明显为媒体资源的请求 media_ext_pattern = re.compile(r'\.(jpg|jpeg|png|gif|bmp|svg|webp|mp4|avi|mov|mp3|wav|flac|woff|woff2|ttf|otf)$', re.IGNORECASE) if media_ext_pattern.search(flow.request.path): flow.response = http.Response.make(404, b"Rejected") return
脚本逻辑说明
- Accept头解析:客户端发起请求时会通过
Accept头告知服务器可接受的内容类型,解析该头直接拦截目标媒体类型的请求,无需转发到服务器; - Application类型过滤:仅放行你指定的安全子类型,拦截其他application类请求;
- URL后缀兜底:针对直接访问资源链接(无Accept头)的场景,通过正则匹配常见媒体后缀补充拦截;
- 匹配到规则后直接返回404响应,实现请求阶段的拦截效果。
内容的提问来源于stack exchange,提问作者old_hands
相关产品推荐
相关产品推荐

