Docker中ONNX Runtime多模型顺序加载推理异常问题排查
问题:Docker部署Flask+ONNX Runtime多模型推理失败,仅首个YOLO模型可用
我开发了一个基于Flask的API,用ONNX Runtime运行YOLO和多个分类器模型,模型由PyTorch训练后转成ONNX格式。本地环境运行正常,能顺序加载并推理不同ONNX模型,但部署到Docker后,只有首个加载的YOLO模型能正常推理,后续分类器模型无法启动推理会话。
推理流程
- 在ONNX Runtime中加载YOLO模型进行初始推理
- 根据YOLO输出裁剪图像
- 将裁剪后的图像依次传入多个ONNX格式的分类器模型
模型输入要求
- YOLO模型:输入图像尺寸为640x640像素
- 分类器模型:输入图像尺寸为224x224像素
怀疑是Docker环境下的资源分配或会话管理问题,想知道能不能通过Docker容器内实现多线程解决,以及具体实现方案。
相关代码片段
Flask路由定义
# Flask app initialization and route definition # ... @app.route("/predict", methods=["POST"]) def predict(): # ... # Step 1: YOLO model to detect boxes yolo_response = yolo_predict(image_np) # ... for box in boxes: # Sequential processing of classifiers stonetype_result = stonetype_predict(resized_image) cut_result = cut_predict(resized_image) color_result = color_predict(resized_image) # ... return jsonify(results) # ...
模型加载与推理函数
def load_model(onnx_file_path): """Load the ONNX model.""" session = ort.InferenceSession(onnx_file_path) return session def infer(session, image_tensor): """Run model inference.""" input_name = session.get_inputs()[0].name output = session.run(None, {input_name: image_tensor}) return output
Docker构建脚本
# Use an official Python runtime as a parent image FROM python:3.10-slim # Set the working directory in the container WORKDIR /usr/src/app # Copy the current directory contents into the container at /usr/src/app COPY . /usr/src/app # Install any needed packages specified in requirements.txt RUN pip install -r requirements.txt # Make port 5000 available to the world outside this container EXPOSE 5000 # Define environment variable ENV MODEL_PATH /usr/src/app/services/yolo_service/yolo.onnx # Run server.py when the container launches CMD ["python", "server.py"]
错误信息
[2024-01-29 13:03:20,899] ERROR in app: Exception on /predict [POST] Traceback (most recent call last): File "/usr/local/lib/python3.10/site-packages/flask/app.py", line 1463, in wsgi_app response = self.full_dispatch_request() File "/usr/local/lib/python3.10/site-packages/flask/app.py", line 872, in full_dispatch_request rv = self.handle_user_exception(e) File "/usr/local/lib/python3.10/site-packages/flask/app.py", line 870, in full_dispatch_request rv = self.dispatch_request() File "/usr/local/lib/python3.10/site-packages/flask/app.py", line 855, in dispatch_request return self.ensure_sync(self.view_functions[rule.endpoint])(**view_args) # type: ignore[no-any-return] File "/usr/src/app/server.py", line 38, in predict stonetype_result = stonetype_predict(resized_image) File "/usr/src/app/services/stonetype_service/app/server.py", line 35, in predict output = sessionStoneType.run(None, {input_name: image_tensor}) File "/usr/local/lib/python3.10/site-packages/onnxruntime/capi/onnxruntime_inference_collection.py", line 220, in run return self._sess.run(output_names, input_feed, run_options) onnxruntime.capi.onnxruntime_pybind11_state.InvalidArgument: [ONNXRuntimeError] : 2 : INVALID_ARGUMENT : Got invalid dimensions for input: images for the following indices index: 2 Got: 224 Expected: 640 index: 3 Got: 224 Expected: 640 Please fix either the inputs or the model.
解决方案
先排查核心问题:模型会话复用错误
从错误信息看,分类器模型接收到224x224的输入却要求640x640,说明分类器的会话实例被错误复用了YOLO模型的会话,或是Docker中模型加载路径混乱,导致分类器加载的实际是YOLO模型。
先做2项验证:
- 进入Docker容器,手动检查
stonetype.onnx等分类器模型的路径和正确性,确认模型输入尺寸为224x224 - 在
load_model函数中添加日志,打印加载模型的输入尺寸:
def load_model(onnx_file_path): session = ort.InferenceSession(onnx_file_path) input_shape = session.get_inputs()[0].shape print(f"Loaded model {onnx_file_path}, input shape: {input_shape}") return session
启动容器后观察日志,确认每个模型的输入尺寸符合预期。
多线程实现方案(解决并发与推理效率问题)
若模型加载无问题,可通过多线程优化推理并发能力:
1. 启用Flask多线程模式
修改server.py的启动逻辑,开启多线程处理请求:
if __name__ == "__main__": app.run(host="0.0.0.0", port=5000, threaded=True)
或直接修改Docker启动命令:
CMD ["python", "-c", "from server import app; app.run(host='0.0.0.0', port=5000, threaded=True)"]
2. 全局预加载模型会话
ONNX Runtime的InferenceSession线程安全,建议Flask启动时全局加载所有模型,避免每次请求重复加载:
# 全局加载所有模型 yolo_session = load_model("services/yolo_service/yolo.onnx") stonetype_session = load_model("services/stonetype_service/stonetype.onnx") cut_session = load_model("services/cut_service/cut.onnx") color_session = load_model("services/color_service/color.onnx") # 推理函数使用全局会话 def yolo_predict(image): tensor = preprocess_yolo(image) # 预处理为640x640 return infer(yolo_session, tensor) def stonetype_predict(image): tensor = preprocess_classifier(image) # 预处理为224x224 return infer(stonetype_session, tensor)
3. 并行推理多个分类器
用ThreadPoolExecutor并行处理多个分类器,提升推理速度:
from concurrent.futures import ThreadPoolExecutor # 全局创建线程池(根据CPU核心数设置max_workers) executor = ThreadPoolExecutor(max_workers=3) @app.route("/predict", methods=["POST"]) def predict(): # ... YOLO推理与图像裁剪 ... for box in boxes: # 并行提交分类器推理任务 futures = [ executor.submit(stonetype_predict, resized_image), executor.submit(cut_predict, resized_image), executor.submit(color_predict, resized_image) ] # 获取并行执行结果 stonetype_result = futures[0].result() cut_result = futures[1].result() color_result = futures[2].result() # ... 结果处理 ... return jsonify(results)
Docker资源优化
若存在资源瓶颈,启动容器时指定CPU和内存配额:
docker run -d --name flask-ml-api -p 5000:5000 --cpus=4 --memory=8g your-image-name
同时优化ONNX Runtime的线程配置:
def load_model(onnx_file_path): session_options = ort.SessionOptions() session_options.intra_op_num_threads = 2 # 算子内并行线程数 session_options.inter_op_num_threads = 2 # 算子间并行线程数 session = ort.InferenceSession(onnx_file_path, sess_options=session_options, providers=["CPUExecutionProvider"]) return session
内容的提问来源于stack exchange,提问作者wadie el
相关产品推荐
相关产品推荐

