You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

API Gateway Lambda代理集成:gzip压缩后JSON出现反斜杠问题

问题:API Gateway集成Lambda启用gzip压缩后出现多余反斜杠

我有一个配置了Lambda代理集成的API Gateway,在Lambda侧启用gzip_b64encode来压缩从S3获取的payload。添加该功能前,输出没有反斜杠,但添加后反斜杠再次出现。

输出示例:

[{"Exchange\":\"TSE\",\"Ticker\":1301,\"Name\":\"Kyokuyo\",\"Country\":\"Japan\",\"Currency\":\"JPY\",\"Asset_Type\":\"Common Stock\",\"ISIN\":\"JP3257200000\",\"Date\":\"4/30/2023\",\"Open\":3580.0,\"High\":3590.0,\"Low\":3545.0,\"Close\":3560.0,\"Volume\":20800,

Lambda代码:

import json
from io import BytesIO
import base64
import gzip
import awswrangler as wr
import boto3

def gzip_b64encode(data):
    compressed = BytesIO()
    with gzip.GzipFile(fileobj=compressed, mode='w') as f:
        json_response = json.dumps(data)
        f.write(json_response.encode('utf-8'))
    return base64.b64encode(compressed.getvalue()).decode('ascii')

def lambda_handler(event, context):
    
    print(event)
    s3 = boto3.client('s3')
    bucket_name = 'tse-candles'
    
    format = event['queryStringParameters']['fmt']
    day = event['queryStringParameters']['date']
    
    A = '123.csv'
    full_path = f"s3://{bucket_name}"

    if day == 'today':
     
      raw_df = wr.s3.read_csv(path=f'{full_path}/{A}', use_threads=True)
      
      if format == 'json':
        
        df = raw_df.to_json(orient="records")
        parsed = json.loads(df)
        parsed_w_separators = json.dumps(parsed, separators=(',', ':'))
      
        return {
          "isBase64Encoded": True,
          "statusCode": 200,
          "headers": { 
            "Content-Type": "application/json",
            'Content-Encoding': 'gzip'
            },
          'body': gzip_b64encode(parsed_w_separators)
          }

原因与解决方案

原因

问题不是gzip编码本身导致的,而是重复JSON序列化造成的:

  • 你已经通过parsed_w_separators = json.dumps(parsed, separators=(',', ':'))生成了最终的JSON字符串(第一次序列化)
  • 但在gzip_b64encode函数里,又对这个字符串执行了json.dumps(data),相当于把JSON字符串当成普通字符串再次序列化,导致所有双引号被转义,出现多余反斜杠

解决方案

方案1:修改压缩函数,直接使用已序列化的JSON字符串

去掉gzip_b64encode里多余的json.dumps,直接写入传入的JSON字符串字节:

def gzip_b64encode(data):
    compressed = BytesIO()
    with gzip.GzipFile(fileobj=compressed, mode='w') as f:
        # 直接写入已序列化的JSON字符串,无需二次序列化
        f.write(data.encode('utf-8'))
    return base64.b64encode(compressed.getvalue()).decode('ascii')

方案2:传入Python对象而非JSON字符串(更推荐)

直接把解析后的Python对象parsed传入压缩函数,让函数内部完成一次序列化:

# lambda_handler中修改调用方式
'body': gzip_b64encode(parsed)

同时保持压缩函数的序列化逻辑(可保留分隔符优化):

def gzip_b64encode(data):
    compressed = BytesIO()
    with gzip.GzipFile(fileobj=compressed, mode='w') as f:
        json_response = json.dumps(data, separators=(',', ':'))
        f.write(json_response.encode('utf-8'))
    return base64.b64encode(compressed.getvalue()).decode('ascii')

额外优化

  • 避免不必要的序列化/反序列化:raw_df.to_json(orient="records")已经生成了JSON字符串,无需再json.loads后json.dumps,直接传入压缩函数可节省性能
  • 当前API Gateway的Content-Encoding和isBase64Encoded配置正确,无需调整

内容的提问来源于stack exchange,提问作者bob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 14:02:46