You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

哪些SageMaker服务器支持server-side batching?如何启用该功能?

在SageMaker端点启用服务器端批处理(MMS、TFServing、TorchServe)

Model Server for Apache MXNet (MMS)

  • 编写model-server.properties配置文件,添加核心批处理参数:
    • max_batch_size: 设置单批可处理的最大请求数量
    • batch_delay: 设置等待凑批的最长延迟时间(单位:毫秒)
    • enable_batch_prediction=true: 显式开启批处理功能
  • 将该配置文件放入模型归档包(.model)的/opt/ml/model目录中
  • 创建SageMaker模型时选择MMS容器,确保启动命令加载该配置文件

TensorFlow Serving (TFServing)

  • 在模型根目录下创建config.pbtxt,配置批处理规则:
    model_config_list {
      config {
        name: "你的模型名称"
        base_path: "/opt/ml/model"
        model_platform: "tensorflow"
        batching_parameters {
          max_batch_size: 32
          batch_timeout_micros: 100000  # 对应100毫秒
          max_enqueued_batches: 10
        }
      }
    }
    
  • 使用TFServing官方容器创建SageMaker端点,容器会自动读取该配置启用批处理

TorchServe

  • 创建config.properties配置文件,添加批处理相关参数:
    inference_address=http://0.0.0.0:8080
    management_address=http://0.0.0.0:8081
    enable_envvars_config=true
    max_batch_size=16
    batch_delay=50  # 50毫秒凑批延迟
    response_timeout=60
    
  • 将该配置文件与模型文件、model-store目录一同打包
  • 在SageMaker中使用TorchServe容器,启动时指定加载该配置;也可直接通过环境变量覆盖参数,例如设置SAGEMAKER_TORCHSERVE_MAX_BATCH_SIZE=16和SAGEMAKER_TORCHSERVE_BATCH_DELAY=50

内容的提问来源于stack exchange,提问作者Francesco Pochetti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 06:15:43