You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scapy嗅探HTTP响应,gzip压缩HTML解压失败求助

解决Scapy嗅探HTTP gzip压缩响应体乱码问题

嘿,我之前也碰到过一模一样的问题!咱们先把事儿捋清楚:你看到的乱码完全是因为响应体被gzip压缩了,直接读packet[Raw].load拿到的是压缩后的二进制流,而解压失败大概率是因为zlib或gzip模块的参数没设置对,或者输入数据的处理方式不对。

下面给你两种亲测有效的解压方案,直接就能用:

方法1:用zlib模块(关键是window参数)

gzip格式属于zlib的一种特殊变体,默认的zlib解压参数识别不了它,得手动设置wbits=16+zlib.MAX_WBITS来告诉模块“这是gzip格式的压缩数据”。代码示例:

import zlib
from scapy.all import sniff, Raw, HTTPResponse

def handle_http_response(packet):
    if packet.haslayer(HTTPResponse) and packet.haslayer(Raw):
        # 先把响应头拎出来
        resp_headers = packet[HTTPResponse].fields
        # 检查是否是gzip压缩
        if 'Content-Encoding' in resp_headers and 'gzip' in resp_headers['Content-Encoding']:
            compressed_data = packet[Raw].load
            try:
                # 核心:设置wbits参数适配gzip格式
                decompressed_data = zlib.decompress(compressed_data, wbits=16+zlib.MAX_WBITS)
                # 转成字符串(编码根据实际情况调整,一般是utf-8)
                html_content = decompressed_data.decode('utf-8')
                print("解压后的HTML内容(前500字符):\n", html_content[:500])
            except Exception as e:
                print(f"解压翻车了: {str(e)}")
        else:
            # 非压缩的直接输出就行
            print(packet[Raw].load.decode('utf-8'))

# 启动嗅探,过滤80端口的HTTP响应
sniff(filter="tcp port 80", prn=handle_http_response, store=0)

方法2:用gzip模块(需要包装成字节流)

gzip模块没法直接处理原始字节数据,得先把压缩数据包装成BytesIO流再读取,不然会报错。代码示例:

import gzip
from io import BytesIO
from scapy.all import sniff, Raw, HTTPResponse

def handle_http_response(packet):
    if packet.haslayer(HTTPResponse) and packet.haslayer(Raw):
        resp_headers = packet[HTTPResponse].fields
        if 'Content-Encoding' in resp_headers and 'gzip' in resp_headers['Content-Encoding']:
            compressed_data = packet[Raw].load
            try:
                # 把字节数据转成gzip能识别的流对象
                with gzip.GzipFile(fileobj=BytesIO(compressed_data), mode='rb') as f:
                    decompressed_data = f.read()
                html_content = decompressed_data.decode('utf-8')
                print("解压后的HTML内容(前500字符):\n", html_content[:500])
            except Exception as e:
                print(f"解压翻车了: {str(e)}")
        else:
            print(packet[Raw].load.decode('utf-8'))

sniff(filter="tcp port 80", prn=handle_http_response, store=0)

额外提醒

  • 如果碰到deflate格式的压缩,把zlib的wbits改成-zlib.MAX_WBITS就行,你可以根据响应头的Content-Encoding动态切换参数。
  • 要是网站用了HTTP分块编码(响应头有Transfer-Encoding: chunked),Scapy的Raw.load可能只拿到其中一块,这时候得维护一个字典跟踪每个TCP连接的分块数据,拼完整后再解压。

内容的提问来源于stack exchange,提问作者James

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 09:14:56