如何通过KServe-TritonServer部署XGBoost模型推理服务?部署遇报错
问题描述
已在KServe默认运行时部署XGBoost模型,现希望切换至KServe-TritonServer部署该模型的推理服务。已知KServe文档显示KServe-TritonServer支持TensorFlow、ONNX、PyTorch、TensorRT,而NVIDIA称Triton推理服务器支持XGBoost模型。尝试使用以下配置部署:
k apply -n kserve-test -f - <<EOF apiVersion: "serving.kserve.io/v1beta1" kind: "InferenceService" metadata: name: "digits-classification-xgboost" spec: predictor: model: modelFormat: name: xgboost protocolVersion: v2 storageUri: "s3://.../digits_classification_model" runtime: kserve-tritonserver EOF
部署后出现报错:
Status: Model Status: Last Failure Info: Message: 指定运行时不支持指定框架/版本 Reason: NoSupportingRuntime States: Active Model State: Target Model State: FailedToLoad Transition Status: InvalidSpec Events: Type Reason Age From Message ---- ------ ---- ---- ------- Warning InternalError 65s (x19 over 9m9s) v1beta1Controllers specified runtime kserve-tritonserver does not support specified framework/version
请问是否有可行方案实现该部署需求?
可行解决方案
KServe官方维护的kserve-tritonserver运行时未内置XGBoost的支持校验逻辑,直接指定modelFormat.name: xgboost会触发兼容性报错,可通过以下两种方案解决:
方式一:转换XGBoost模型为ONNX格式
Triton原生支持ONNX模型,通过转换模型格式绕过KServe的运行时校验,操作步骤:
- 转换模型:使用
xgboost和onnxmltools工具完成格式转换,示例代码:
import xgboost as xgb from onnxmltools.convert import convert_xgboost # 加载已训练的XGBoost模型 model = xgb.Booster() model.load_model("digits_classification_model.model") # 转换为ONNX格式 onnx_model = convert_xgboost(model) with open("digits_classification_model.onnx", "wb") as f: f.write(onnx_model.SerializeToString())
- 上传转换后的ONNX模型到存储路径(如原S3路径或新路径),修改部署配置:
k apply -n kserve-test -f - <<EOF apiVersion: "serving.kserve.io/v1beta1" kind: "InferenceService" metadata: name: "digits-classification-xgboost" spec: predictor: model: modelFormat: name: onnx protocolVersion: v2 storageUri: "s3://.../digits_classification_model_onnx" runtime: kserve-tritonserver EOF
方式二:自定义KServe-TritonServer运行时镜像
若需保留XGBoost原生格式,可自定义包含XGBoost后端的Triton镜像,操作步骤:
- 构建自定义镜像:基于官方Triton镜像添加XGBoost后端,Dockerfile示例:
FROM nvcr.io/nvidia/tritonserver:23.10-py3 # 安装XGBoost依赖 RUN pip install xgboost==2.0.3 # 复制预编译的XGBoost后端到Triton后端目录(可从Triton官方源码编译获取) COPY xgboost_backend /opt/tritonserver/backends/xgboost
- 推送自定义镜像到私有仓库后,修改InferenceService配置指定该镜像:
k apply -n kserve-test -f - <<EOF apiVersion: "serving.kserve.io/v1beta1" kind: "InferenceService" metadata: name: "digits-classification-xgboost" spec: predictor: model: modelFormat: name: xgboost protocolVersion: v2 storageUri: "s3://.../digits_classification_model" runtime: kserve-tritonserver container: image: your-custom-triton-image:tag EOF
- 按Triton要求组织模型目录结构:
digits_classification_model/ ├── 1/ │ └── model.xgb └── config.pbtxt
config.pbtxt配置示例:
name: "digits_classification_xgboost" platform: "xgboost" max_batch_size: 32 input [ { name: "input" data_type: TYPE_FP32 dims: [784] } ] output [ { name: "output" data_type: TYPE_FP32 dims: [10] } ]
内容的提问来源于stack exchange,提问作者HoonCheol Shin
相关产品推荐
相关产品推荐

