You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Python字节码解码为ASCII?Selenium获取XML网络响应遇阻

Selenium提取网络响应无法解析为XML,得到乱码字节数据

问题详情

使用Selenium抓取网络响应时,预期获取XML格式内容,但实际得到的是乱码字节数据:
['b\'\xa5\xff\xff\xc7\x88\xe4\xb4\xd7\x03\xa0\x11:|\xce\xdb\xb7\x0f\xf1\xdf\xfc\x1f\xdb\x93\x91^\xbc\xa3\xdd\xc2\x02V\x00\xba$\xbd\x10\xd2\xd0E\xf2\x90\xb6\xca\xee\x10\xbf\xbf_\xbf\xfc\xef?\xe9\x13{H\xf1\xa1\xa0\x00\x1c\x01(\x80\x1c\x81\x02(s\xe7Z\xf3\xb3N\xf5L\xdc>\xe7\x8f\xbbwl\xbf\x99\x91\xd4O\xde\xb4,\xf3PH\x02L1\x00\xc98\xc3,\x13!\x82\xc6\xc2\xa6Bd"k\xcb\x9d(\xb9\x13%WQr\x15%W\xb1\xe5J\t\x9e:\x8a\x03\x99\x06H\xd0\x8f\xd8\xfe\x9f9\xbc\xfc\x157\x111\xd7\x15\xaab\xfb\xe8;\xab\xee\xfc\x9b\xeeu\x10<d\x04\x06Y\xa8\xd7\x9f\x11...']

当前使用的代码片段:

for request in driver.requests:
    if request.response:
        text_file.write(str(request.response.body))

已尝试的方法及错误

  • 直接用ASCII/UTF-8/CP1251/CP1252解码字节数据:

    decoded = request.response.body.decode('ascii')
    # 或
    decoded = request.response.body.decode('utf-8')
    

    均触发解码错误:
    UnicodeDecodeError: 'ascii' codec can't decode byte 0xa5 in position 0: ordinal not in range(128)

  • 尝试Base64解码:

    decoded = base64.b64decode(request.response.body)
    

    得到的结果仍是乱码字节:b'T@\x00\xad\x9a\xb5\xba\xfa3u\xca\x84PG\xbd\x8a\xab\x1f\xcdcJ%\r\xd4\xff\x0c$)\x9a>....,结合解码后转ASCII也无法解决,同样报UnicodeDecodeError。

预期响应为约1.5MB的XML内容。

解决方案

这种情况大概率是服务器返回的响应被压缩(gzip/deflate格式),Selenium直接返回了压缩后的原始字节,需要先解压再解码:

  1. 先检查响应头的Content-Encoding字段,确认压缩类型
  2. 根据压缩类型解压字节,再用UTF-8(XML常用编码)解码为文本

示例代码:

import gzip
import zlib

for request in driver.requests:
    if request.response:
        raw_body = request.response.body
        # 获取响应头中的压缩编码信息
        content_encoding = request.response.headers.get('Content-Encoding', '').lower()
        
        # 解压数据
        if 'gzip' in content_encoding:
            decompressed_body = gzip.decompress(raw_body)
        elif 'deflate' in content_encoding:
            decompressed_body = zlib.decompress(raw_body)
        else:
            # 如果没有压缩标识,也可以尝试强制用gzip解压(部分服务器可能未正确设置头)
            try:
                decompressed_body = gzip.decompress(raw_body)
            except:
                decompressed_body = raw_body
        
        # 解码为XML文本(优先用UTF-8,若不行可尝试其他编码如GBK)
        try:
            xml_content = decompressed_body.decode('utf-8')
        except UnicodeDecodeError:
            xml_content = decompressed_body.decode('gbk')
        
        text_file.write(xml_content)

如果上述方法仍不生效,可以通过浏览器开发者工具查看该请求的响应头和实际响应内容,确认是否有其他编码或加密方式(比如自定义压缩)。


内容的提问来源于stack exchange,提问作者phdBrahmana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 05:45:34