Python requests调用API返回无效数据,仅虚拟环境中出现异常
FIPE车辆API JSON解析异常问题
问题现象
调用巴西FIPE车辆API(https://veiculos.fipe.org.br)时,多数请求能返回正常JSON响应,但少数请求返回的内容带有\x8b\x15\x80前缀和\x03后缀的无效字符,直接解析会触发JSONDecodeError。用相同参数在Firefox浏览器中请求可得到正常JSON,且该问题仅在Python虚拟环境(venv)中出现。
测试代码
import requests headers = { 'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:144.0) Gecko/20100101 Firefox/144.0', 'Accept': 'application/json, text/javascript, */*; q=0.01', 'Accept-Language': 'en-US,en;q=0.8,pt-BR;q=0.5,pt;q=0.3', 'Accept-Encoding': 'gzip, deflate, br, zstd', 'Content-Type': 'application/x-www-form-urlencoded; charset=UTF-8', 'X-Requested-With': 'XMLHttpRequest', 'Origin': 'https://veiculos.fipe.org.br', 'DNT': '1', 'Sec-GPC': '1', 'Connection': 'keep-alive', 'Referer': 'https://veiculos.fipe.org.br/', 'Cookie': 'ROUTEID=.5', 'Sec-Fetch-Dest': 'empty', 'Sec-Fetch-Mode': 'cors', 'Sec-Fetch-Site': 'same-origin', 'Priority': 'u=0', 'TE': 'trailers', } def api_call(url, headers, parameters): try: r = requests.request('post', url, headers=headers, data=parameters) print(r.content) r_json = r.json() r.close() return r_json except Exception as e: print(repr(e)) url = 'https://veiculos.fipe.org.br/api/veiculos/ConsultarAnoModelo' # 正常返回的请求 parameters = [('codigoModelo', 6906), ('codigoTabelaReferencia', 327), ('codigoTipoVeiculo', 1), ('CodigoMarca', 189)] r_json = api_call(url, headers, parameters) # 返回无效字符的请求 parameters = [('codigoModelo', 6340), ('codigoTabelaReferencia', 327), ('codigoTipoVeiculo', 1), ('CodigoMarca', 189)] r_json = api_call(url, headers, parameters)
错误输出
b'[{"Label":"2016 Gasolina","Value":"2016-1"},{"Label":"2014 Gasolina","Value":"2014-1"}]' b'\x8b\x15\x80[{"Label":"2011 Gasolina","Value":"2011-1"}]\x03' JSONDecodeError('Expecting value: line 1 column 1 (char 0)')
原因分析
这些无效前缀和后缀是不完整的gzip压缩残留数据。虽然requests库默认会自动处理gzip压缩响应,但少数情况下服务器返回的压缩数据格式不规范,导致requests的自动解压逻辑未完全清理压缩头/尾。另外,虚拟环境中requests或其依赖库(如urllib3)的版本与全局环境存在差异,可能是触发该兼容性问题的原因。
解决方案
1. 手动清理无效字符后解析
直接定位JSON内容的起始和结束位置,过滤掉前后无效字符:
import json def api_call(url, headers, parameters): try: r = requests.request('post', url, headers=headers, data=parameters) content = r.content # 找到JSON起始符号({或[)的索引 start_idx = next(i for i, c in enumerate(content) if c in b'{[') # 找到JSON结束符号(}或])的索引 end_idx = len(content) - 1 - next(i for i, c in enumerate(reversed(content)) if c in b'}]') # 截取有效JSON部分 cleaned_content = content[start_idx:end_idx+1] r_json = json.loads(cleaned_content) r.close() return r_json except Exception as e: print(repr(e))
2. 禁用自动压缩或手动处理解压
可以通过两种方式避免不规范的压缩问题:
- 移除压缩请求头:让服务器返回未压缩的原始JSON数据
# 从请求头中删除Accept-Encoding字段 headers.pop('Accept-Encoding', None)
- 手动解压响应:尝试手动gzip解压,失败则回退到清理无效字符
import json import gzip from io import BytesIO def api_call(url, headers, parameters): try: r = requests.request('post', url, headers=headers, data=parameters) try: # 尝试手动解压gzip数据 content = gzip.decompress(r.content) except (gzip.BadGzipFile, OSError): # 解压失败则清理无效字符 content = r.content start_idx = next(i for i, c in enumerate(content) if c in b'{[') end_idx = len(content) - 1 - next(i for i, c in enumerate(reversed(content)) if c in b'}]') content = content[start_idx:end_idx+1] r_json = json.loads(content) r.close() return r_json except Exception as e: print(repr(e))
3. 更新虚拟环境依赖库
确保虚拟环境中的requests和urllib3为最新版本,修复可能存在的解压逻辑bug:
pip install --upgrade requests urllib3
内容的提问来源于stack exchange,提问作者Clodoaldo Pinto
相关产品推荐
相关产品推荐

