在Google Vertex AI部署带自定义Handler的TorchServe实例时的返回格式兼容问题求助
在Google Vertex AI部署带自定义Handler的TorchServe实例解决方案
我之前也遇到过完全一样的问题,折腾了好一阵子才找到关键所在——问题出在TorchServe对postprocess方法的返回值要求,和Vertex AI的响应格式要求之间的冲突。下面直接给你可行的解决思路和代码示例:
核心问题解析
- TorchServe的
postprocess方法必须返回一个列表,列表中的每个元素对应一个输入样本的预测结果(比如你的[[1, 2, 1], [2, 3, 3]]就是两个样本的嵌入向量)。如果返回字典,TorchServe会判定输出无效,抛出Invalid model predict output错误。 - Vertex AI要求响应必须是
{"predictions": PREDICTIONS}格式的JSON对象,其中PREDICTIONS是预测结果数组。直接返回列表的话,Vertex AI无法识别格式,会触发ModelNotFoundException。
正确的实现方式
不要在postprocess里包装字典,而是重写Handler的handle方法,在最后把postprocess返回的列表包装成符合Vertex AI要求的格式。
完整的自定义Handler代码示例
from ts.torch_handler.base_handler import BaseHandler class EmbeddingHandler(BaseHandler): def preprocess(self, data): # 这里写你的预处理逻辑,比如解析输入文本、转换为模型需要的张量 processed_input = [] for item in data: # 示例:假设输入是{"text": "sample sentence"}格式 text = item.get("text") # 替换成你的预处理步骤(比如tokenize) processed_input.append(your_tokenizer(text)) return processed_input def inference(self, data): # 执行模型推理 with torch.no_grad(): outputs = self.model(*data) return outputs def postprocess(self, data): # 这里仍然返回预测结果的列表(保持TorchServe要求的格式) # 示例:把张量转换为列表格式 return [output.tolist() for output in data] def handle(self, data, context): # 重写handle方法,在最后包装成Vertex AI需要的格式 input_batch = self.preprocess(data) output_batch = self.inference(input_batch) predictions = self.postprocess(output_batch) # 包装成符合Vertex AI要求的字典 return {"predictions": predictions}
验证步骤
- 本地测试:启动TorchServe后,发送请求到预测端点,确认返回的响应是
{"predictions": [[1,2,1], [2,3,3]]}格式:curl -X POST http://localhost:8080/predictions/your_model -H "Content-Type: application/json" -d '[{"text": "test sentence 1"}, {"text": "test sentence 2"}]' - 部署到Vertex AI:用修改后的Handler重新打包模型镜像,部署后测试预测请求,应该就能正常返回结果了。
额外注意事项
- 确保容器暴露的端口是Vertex AI要求的默认端口(8080),启动命令正确(比如
torchserve --start --model-store model_store --models your_model=model.mar)。 - 如果你的模型有特殊的输入格式,要保证
preprocess方法能正确解析Vertex AI发送的请求体(通常是JSON数组格式)。
内容的提问来源于stack exchange,提问作者Timon. Z
相关产品推荐
相关产品推荐

