API Gateway Lambda代理集成:gzip压缩后JSON出现反斜杠问题
问题:API Gateway集成Lambda启用gzip压缩后出现多余反斜杠
我有一个配置了Lambda代理集成的API Gateway,在Lambda侧启用gzip_b64encode来压缩从S3获取的payload。添加该功能前,输出没有反斜杠,但添加后反斜杠再次出现。
输出示例:
[{"Exchange\":\"TSE\",\"Ticker\":1301,\"Name\":\"Kyokuyo\",\"Country\":\"Japan\",\"Currency\":\"JPY\",\"Asset_Type\":\"Common Stock\",\"ISIN\":\"JP3257200000\",\"Date\":\"4/30/2023\",\"Open\":3580.0,\"High\":3590.0,\"Low\":3545.0,\"Close\":3560.0,\"Volume\":20800,
Lambda代码:
import json from io import BytesIO import base64 import gzip import awswrangler as wr import boto3 def gzip_b64encode(data): compressed = BytesIO() with gzip.GzipFile(fileobj=compressed, mode='w') as f: json_response = json.dumps(data) f.write(json_response.encode('utf-8')) return base64.b64encode(compressed.getvalue()).decode('ascii') def lambda_handler(event, context): print(event) s3 = boto3.client('s3') bucket_name = 'tse-candles' format = event['queryStringParameters']['fmt'] day = event['queryStringParameters']['date'] A = '123.csv' full_path = f"s3://{bucket_name}" if day == 'today': raw_df = wr.s3.read_csv(path=f'{full_path}/{A}', use_threads=True) if format == 'json': df = raw_df.to_json(orient="records") parsed = json.loads(df) parsed_w_separators = json.dumps(parsed, separators=(',', ':')) return { "isBase64Encoded": True, "statusCode": 200, "headers": { "Content-Type": "application/json", 'Content-Encoding': 'gzip' }, 'body': gzip_b64encode(parsed_w_separators) }
原因与解决方案
原因
问题不是gzip编码本身导致的,而是重复JSON序列化造成的:
- 你已经通过
parsed_w_separators = json.dumps(parsed, separators=(',', ':'))生成了最终的JSON字符串(第一次序列化) - 但在
gzip_b64encode函数里,又对这个字符串执行了json.dumps(data),相当于把JSON字符串当成普通字符串再次序列化,导致所有双引号被转义,出现多余反斜杠
解决方案
方案1:修改压缩函数,直接使用已序列化的JSON字符串
去掉gzip_b64encode里多余的json.dumps,直接写入传入的JSON字符串字节:
def gzip_b64encode(data): compressed = BytesIO() with gzip.GzipFile(fileobj=compressed, mode='w') as f: # 直接写入已序列化的JSON字符串,无需二次序列化 f.write(data.encode('utf-8')) return base64.b64encode(compressed.getvalue()).decode('ascii')
方案2:传入Python对象而非JSON字符串(更推荐)
直接把解析后的Python对象parsed传入压缩函数,让函数内部完成一次序列化:
# lambda_handler中修改调用方式 'body': gzip_b64encode(parsed)
同时保持压缩函数的序列化逻辑(可保留分隔符优化):
def gzip_b64encode(data): compressed = BytesIO() with gzip.GzipFile(fileobj=compressed, mode='w') as f: json_response = json.dumps(data, separators=(',', ':')) f.write(json_response.encode('utf-8')) return base64.b64encode(compressed.getvalue()).decode('ascii')
额外优化
- 避免不必要的序列化/反序列化:
raw_df.to_json(orient="records")已经生成了JSON字符串,无需再json.loads后json.dumps,直接传入压缩函数可节省性能 - 当前API Gateway的
Content-Encoding和isBase64Encoded配置正确,无需调整
内容的提问来源于stack exchange,提问作者bob
相关产品推荐
相关产品推荐

