You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在AWS SageMaker多模型端点传递tritonclient.http推理请求及适配单输入

问题

我们基于AWS SageMaker搭建了搭载NVIDIA Triton Server的多模型端点,使用tritonclient.http的httpclient.InferenceServerClient.generate_request_body方法构建推理请求载荷。目前看到的示例里inputs和outputs都是传入列表形式,想知道有没有仅传入单个输入的示例?另外,后端处理请求的model.py当前只支持处理多输入列表,能不能修改成支持单输入?

原请求构建代码:

import tritonclient.http as httpclient
import numpy as np


def get_text_payload_binary(model_name, text):
    inputs = []
    outputs = []
    input_ids, attention_mask = tokenize_text(model_name, text)
    inputs.append(httpclient.InferInput("input_ids", input_ids.shape, "INT32"))
    inputs.append(httpclient.InferInput("attention_mask", attention_mask.shape, "INT32"))

    inputs[0].set_data_from_numpy(input_ids.astype(np.int32), binary_data=True)
    inputs[1].set_data_from_numpy(attention_mask.astype(np.int32), binary_data=True)

    output_name = "output" if model_name == "t5-small" else "logits"
    request_body, header_length = httpclient.InferenceServerClient.generate_request_body(
        inputs, outputs=outputs
    )
    return request_body, header_length

原model.py代码:

import numpy as np
import sys
import os
import json
from pathlib import Path

import torch

import triton_python_backend_utils as pb_utils

class TritonPythonModel:
  
    def initialize(self, args):
         ...


    def execute(self, requests):
        """`execute` must be implemented in every Python model. `execute`
        function receives a list of pb_utils.InferenceRequest as the only
        argument. This function is called when an inference is requested
        for this model.
        Parameters
        ----------
        requests : list
          A list of pb_utils.InferenceRequest
        Returns
        -------
        list
          A list of pb_utils.InferenceResponse. The length of this list must
          be the same as `requests`
        """
        responses = []
        for request in requests:
            input_ids = pb_utils.get_input_tensor_by_name(request, "input_ids")
            input_ids = input_ids.as_numpy()
            input_ids = torch.as_tensor(input_ids).long().cuda()
            attention_mask = pb_utils.get_input_tensor_by_name(request, "attention_mask")
            attention_mask = attention_mask.as_numpy()
            attention_mask = torch.as_tensor(attention_mask).long().cuda()
            inputs = {'input_ids': input_ids, 'attention_mask': attention_mask}
            translation = self.model.generate(**inputs, num_beams=1)
           
            np_translation =  translation.cpu().int().detach().numpy()
            inference_response = pb_utils.InferenceResponse(
                output_tensors=[
                    pb_utils.Tensor(
                        "output",
                        np_translation.astype(self.output_dtype)
                    )
                ]
            )
            responses.append(inference_response)
        return responses

解答

一、单个输入的请求构建示例

完全支持传入单个输入,只需构造仅包含一个InferInput对象的列表即可。以下是针对仅传input_ids的修改示例:

import tritonclient.http as httpclient
import numpy as np

def get_single_input_payload_binary(model_name, text):
    inputs = []
    outputs = []
    # 根据模型需求调整tokenize逻辑,这里假设只返回input_ids
    input_ids = tokenize_text(model_name, text)
    # 仅添加一个输入张量
    inputs.append(httpclient.InferInput("input_ids", input_ids.shape, "INT32"))
    inputs[0].set_data_from_numpy(input_ids.astype(np.int32), binary_data=True)

    output_name = "output" if model_name == "t5-small" else "logits"
    request_body, header_length = httpclient.InferenceServerClient.generate_request_body(
        inputs, outputs=outputs
    )
    return request_body, header_length

二、修改model.py支持单输入

要兼容单输入场景,核心是避免硬编码依赖所有输入张量,改为动态检测请求中存在的输入。修改后的execute方法如下:

def execute(self, requests):
    responses = []
    for request in requests:
        inputs = {}
        
        # 获取必填的input_ids
        input_ids = pb_utils.get_input_tensor_by_name(request, "input_ids")
        input_ids = torch.as_tensor(input_ids.as_numpy()).long().cuda()
        inputs['input_ids'] = input_ids
        
        # 尝试获取可选的attention_mask,不存在则生成默认值
        try:
            attention_mask = pb_utils.get_input_tensor_by_name(request, "attention_mask")
            if attention_mask is not None:
                attention_mask = torch.as_tensor(attention_mask.as_numpy()).long().cuda()
                inputs['attention_mask'] = attention_mask
        except pb_utils.TritonModelException:
            # 生成与input_ids形状一致的全1attention_mask作为默认值
            attention_mask = torch.ones_like(input_ids).long().cuda()
            inputs['attention_mask'] = attention_mask
        
        # 动态传入所有存在的输入
        translation = self.model.generate(**inputs, num_beams=1)
        
        np_translation = translation.cpu().int().detach().numpy()
        inference_response = pb_utils.InferenceResponse(
            output_tensors=[
                pb_utils.Tensor(
                    "output",
                    np_translation.astype(self.output_dtype)
                )
            ]
        )
        responses.append(inference_response)
    return responses

关键调整说明

  • 使用try-except捕获attention_mask不存在的异常,避免请求直接失败
  • 若模型必须依赖attention_mask,自动生成全1的默认张量(与input_ids形状匹配)
  • 通过字典动态收集输入,使用**inputs解包传入模型,同时兼容单输入和多输入请求

内容的提问来源于stack exchange,提问作者haju

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 22:23:14