You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SageMaker Serverless推理请求体过大问题及体积优化问询

问题描述

我在S3存储了一个图像分类模型,部署为实时推理端点时可处理任意尺寸的图像,但部署为Serverless推理端点时,无法处理尺寸大于400x400的图像,报错信息:

ValidationError: An error occurred (ValidationError) when calling the InvokeEndpoint operation: Request {request-id} has oversized body.

复现步骤

1. 部署Serverless模型

from sagemaker.tensorflow import TensorFlowModel
from sagemaker import get_execution_role
from sagemaker import Session
import boto3
from sagemaker.serverless import ServerlessInferenceConfig
print('starting ...')

model_data = "s3://datascience--sagemaker/model_repository/Reimbursement Flow/blur_classifier/model.tar.gz"

role = get_execution_role()
sess = Session()
bucket = sess.default_bucket()
region = boto3.Session().region_name

tf_framework_version = '2.0.0'

sm_model = TensorFlowModel(model_data = model_data,
                 framework_version = tf_framework_version,
                 role=role)

predictor = sm_model.deploy(
    endpoint_name = 'blur-classifier-serverless',
    serverless_inference_config = ServerlessInferenceConfig(
                memory_size_in_mb= 2048,
                max_concurrency= 1,
            )
)

2. 执行预测

from PIL import Image 
import numpy as np 
import boto3 
runtime = boto3.client("sagemaker-runtime") 
import json
img_file="doc_classifier_images/012_0205.jpg"
img_file="0ed2d3e1-fe50-4026-b3bf-9aef535f48cc.jpg"
img = Image.open(img_file)
size = 600
img = img.resize((size, size))
img = np.array(img)
img = img.reshape((1, size, size, 3))
img = img/255.
img = np.around(img, decimals=3)
payload = json.dumps(np.asarray(img).astype(float).tolist())

model_name = "blur-classifier-serverless"
content_type = "application/json" 
response = runtime.invoke_endpoint( EndpointName=model_name, ContentType=content_type, Body=payload) 
pred=json.load(response['Body'])

预期行为

应成功完成预测。设置size = 400时可正常工作,部署为实时推理端点时,400和600尺寸的图像均能正常处理。

补充信息

实时推理模型部署脚本

from sagemaker.tensorflow import TensorFlowModel
from sagemaker import get_execution_role
from sagemaker import Session
print('starting ...')

model_data = "s3://datascience--sagemaker/model_repository/Reimbursement Flow/blur_classifier/model.tar.gz"
instance_type = "ml.m4.xlarge"

role = get_execution_role()
sess = Session()
bucket = sess.default_bucket()

instance_type = 'ml.m4.xlarge'
tf_framework_version = '2.0.0'
temp_endpoint_name = "temp"

sm_model = TensorFlowModel(model_data = model_data,
                 framework_version = tf_framework_version,
                 role=role)

# Now to deploy the model
tf_predictor = sm_model.deploy(endpoint_name="blurclassifier-server",
                               initial_instance_count=1,
                               instance_type=instance_type,
                              )

请求体大小分析

请求体大小超过6MB,各步骤对象大小如下:

size = 600
img = img.resize((size, size))
print('2', sys.getsizeof(img))

img = np.array(img)
print('3', sys.getsizeof(img))

img = img.reshape((1, size, size, 3))
print('4', sys.getsizeof(img))

img = img / 255.0
print('5', sys.getsizeof(img))

img = np.around(img, decimals=3)
print('6', sys.getsizeof(img))

payload = json.dumps(img.tolist())
byte_ = payload.encode("utf-8")
size_in_bytes = len(byte_)
print('7', size_in_bytes)

输出:

1 48
2 48
3 1080144
4 160
5 8640160
6 8640160
7 8136202

解决方案

1. 改用二进制格式传输图像(最有效)

JSON格式会大幅增加数据体积,直接传输图像的二进制字节是最小体积的方式。修改预测代码:

from PIL import Image 
import boto3 
import io

runtime = boto3.client("sagemaker-runtime") 
img_file="0ed2d3e1-fe50-4026-b3bf-9aef535f48cc.jpg"
img = Image.open(img_file)
size = 600
img = img.resize((size, size))

# 将图像转为二进制字节流
buffer = io.BytesIO()
img.save(buffer, format='JPEG')  # 根据图像类型选JPEG/PNG
buffer.seek(0)
payload = buffer.read()

model_name = "blur-classifier-serverless"
content_type = "image/jpeg"  # 对应保存的格式
response = runtime.invoke_endpoint(
    EndpointName=model_name,
    ContentType=content_type,
    Body=payload
) 
pred=json.load(response['Body'])

同时需要修改模型的推理代码(inference.py),使其能处理二进制图像输入:

import tensorflow as tf
import numpy as np
from PIL import Image
import io

def handler(data, context):
    # 解析二进制图像
    img = Image.open(io.BytesIO(data))
    img = img.resize((600, 600))  # 可根据模型需求调整
    img_array = np.array(img) / 255.0
    img_array = np.expand_dims(img_array, axis=0)
    
    # 加载模型并预测
    model = tf.keras.models.load_model('model')
    predictions = model.predict(img_array)
    return predictions.tolist()

2. 压缩JSON payload(如需保留JSON格式)

  • 降低数据精度:保留原始像素值(0-255整数),无需提前归一化,在模型侧处理:
# 修改预测代码
img = np.array(img).astype(np.uint8)  # 保留0-255整数
payload = json.dumps(img.tolist())

模型侧添加归一化步骤:img_array = img_array / 255.0

  • 使用gzip压缩传输:
import gzip
import base64

payload = json.dumps(img.tolist())
compressed_payload = gzip.compress(payload.encode('utf-8'))
# 转为base64字符串避免二进制传输问题
payload_b64 = base64.b64encode(compressed_payload).decode('utf-8')

# 调用端点时指定压缩格式
response = runtime.invoke_endpoint(
    EndpointName=model_name,
    ContentType="application/json",
    ContentEncoding="gzip",
    Body=payload_b64
)

模型侧需先解码解压:

import gzip
import base64
import json

def handler(data, context):
    payload_b64 = data.decode('utf-8')
    compressed_payload = base64.b64decode(payload_b64)
    payload = gzip.decompress(compressed_payload).decode('utf-8')
    img_array = np.array(json.loads(payload))
    # 后续处理逻辑...

3. 调整图像参数

  • 如果业务允许,适当降低图像尺寸(如从600x600降至500x500),直接减少数据量。
  • 将RGB图像转为灰度图(单通道),体积变为原来的1/3:
img = img.convert('L')  # 转为灰度图

内容的提问来源于stack exchange,提问作者sid8491

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 22:15:53