Python中使用httplib2请求URL时内容解码问题求助
解决httplib2获取URL返回二进制压缩数据的问题
问题原因
- GET请求中不必要地设置了
Content-Type头,会干扰服务器的正常响应逻辑 - 请求头声明支持
gzip, deflate, br压缩编码,但httplib2默认不会自动解压这些压缩后的响应内容,返回的是原始二进制压缩数据,而curl会自动完成解压步骤
修改后的代码
from __future__ import unicode_literals import httplib2 import gzip import io import zlib from bs4 import BeautifulSoup def initialize(): global url url = "http://nottherealurl.com" global header header = set_header() def set_header(): return { "Accept":"text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8", "Accept-Encoding":"gzip, deflate, br", "Accept-Language":"en-US,en;q=0.5", "Connection":"keep-alive", "DNT":"1", "Sec-Fetch-Dest":"document", "Sec-Fetch-Mode":"navigate", "Sec-Fetch-Site":"cross-site", "Sec-Fetch-User":"?1", "Sec-GPC":"1", "Upgrade-Insecure-Requests":"1", "TE":"trailers", "User-Agent":"Mozilla/5.0 (Windows NT 10.0; rv:122.0) Gecko/20100101 Firefox/122.0" } def get_url(): initialize() h = httplib2.Http() (resp, content) = h.request(url,"GET",headers=header) # 根据响应头编码格式解压内容 content_encoding = resp.get('content-encoding', '') if content_encoding == 'gzip': buf = io.BytesIO(content) with gzip.GzipFile(fileobj=buf) as f: content = f.read().decode('utf-8') elif content_encoding == 'deflate': content = zlib.decompress(content).decode('utf-8') elif content_encoding == 'br': # 需先安装brotli库:pip install brotli import brotli content = brotli.decompress(content).decode('utf-8') else: content = content.decode('utf-8') print(content)
额外说明
- 移除了GET请求中多余的
Content-Type头,该头仅在POST/PUT等带请求体的请求中需要 - 增加了根据响应的
content-encoding头自动解压内容的逻辑,覆盖三种常见压缩格式 - 若需支持brotli压缩,需先执行
pip install brotli安装对应库 - 移除了原代码中未使用的
requests和subprocess导入,精简代码结构
内容的提问来源于stack exchange,提问作者Z T Minhas
相关产品推荐
相关产品推荐

