You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用Selenium获取网络请求响应体遇字节解码问题求助

Selenium获取HTTP请求响应内容的解码方案

request.response.body返回的是响应体的原始字节流,需要按响应指定的编码规则解码才能得到可读的文本内容,可按以下步骤处理:

  • 提取响应编码
    优先从响应头的Content-Type字段中提取charset参数作为解码编码,无匹配值时用utf-8兜底,示例代码如下:
    import cgi
    # 读取响应头的Content-Type值
    content_type = request.response.headers.get('Content-Type', '')
    # 解析Content-Type中的参数
    _, params = cgi.parse_header(content_type)
    # 提取编码,无则默认utf-8
    encoding = params.get('charset', 'utf-8')
    
  • 解码字节流
    使用提取到的编码对原始字节流解码,增加错误忽略参数避免特殊字符导致解码中断:
    readable_body = request.response.body.decode(encoding, errors='ignore')
    
  • 压缩响应处理
    如果响应头Content-Encoding字段标记内容为gzip/deflate压缩格式,需要先解压再解码:
    import gzip
    import zlib
    
    content_encoding = request.response.headers.get('Content-Encoding', '')
    raw_body = request.response.body
    # 先解压
    if 'gzip' in content_encoding:
        raw_body = gzip.decompress(raw_body)
    elif 'deflate' in content_encoding:
        raw_body = zlib.decompress(raw_body, -zlib.MAX_WBITS)
    # 再解码
    readable_body = raw_body.decode(encoding, errors='ignore')
    
  • 注意事项
    如果使用Selenium 4+的DevTools协议监听网络请求,需要等待responseReceived事件触发完成后再获取响应体,避免拿到不完整的字节流导致解码失败。

内容的提问来源于stack exchange,提问作者Jashan Pabla

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 03:15:07