You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Docker中ONNX Runtime多模型顺序加载推理异常问题排查

问题:Docker部署Flask+ONNX Runtime多模型推理失败,仅首个YOLO模型可用

我开发了一个基于Flask的API,用ONNX Runtime运行YOLO和多个分类器模型,模型由PyTorch训练后转成ONNX格式。本地环境运行正常,能顺序加载并推理不同ONNX模型,但部署到Docker后,只有首个加载的YOLO模型能正常推理,后续分类器模型无法启动推理会话。

推理流程

  1. 在ONNX Runtime中加载YOLO模型进行初始推理
  2. 根据YOLO输出裁剪图像
  3. 将裁剪后的图像依次传入多个ONNX格式的分类器模型

模型输入要求

  • YOLO模型:输入图像尺寸为640x640像素
  • 分类器模型:输入图像尺寸为224x224像素

怀疑是Docker环境下的资源分配或会话管理问题,想知道能不能通过Docker容器内实现多线程解决,以及具体实现方案。


相关代码片段

Flask路由定义

# Flask app initialization and route definition
# ...
@app.route("/predict", methods=["POST"])
def predict():
    # ...
    # Step 1: YOLO model to detect boxes
    yolo_response = yolo_predict(image_np)
    # ...
    for box in boxes:
        # Sequential processing of classifiers
        stonetype_result = stonetype_predict(resized_image)
        cut_result = cut_predict(resized_image)
        color_result = color_predict(resized_image)
        # ...
    return jsonify(results)
# ...

模型加载与推理函数

def load_model(onnx_file_path):
    """Load the ONNX model."""
    session = ort.InferenceSession(onnx_file_path)
    return session

def infer(session, image_tensor):
    """Run model inference."""
    input_name = session.get_inputs()[0].name
    output = session.run(None, {input_name: image_tensor})
    return output

Docker构建脚本

# Use an official Python runtime as a parent image
FROM python:3.10-slim

# Set the working directory in the container
WORKDIR /usr/src/app

# Copy the current directory contents into the container at /usr/src/app
COPY . /usr/src/app

# Install any needed packages specified in requirements.txt
RUN pip install -r requirements.txt

# Make port 5000 available to the world outside this container
EXPOSE 5000

# Define environment variable
ENV MODEL_PATH /usr/src/app/services/yolo_service/yolo.onnx

# Run server.py when the container launches
CMD ["python", "server.py"]

错误信息

[2024-01-29 13:03:20,899] ERROR in app: Exception on /predict [POST]
Traceback (most recent call last):
File "/usr/local/lib/python3.10/site-packages/flask/app.py", line 1463, in wsgi_app
response = self.full_dispatch_request()
File "/usr/local/lib/python3.10/site-packages/flask/app.py", line 872, in full_dispatch_request
rv = self.handle_user_exception(e)
File "/usr/local/lib/python3.10/site-packages/flask/app.py", line 870, in full_dispatch_request
rv = self.dispatch_request()
File "/usr/local/lib/python3.10/site-packages/flask/app.py", line 855, in dispatch_request
return self.ensure_sync(self.view_functions[rule.endpoint])(**view_args) # type: ignore[no-any-return]
File "/usr/src/app/server.py", line 38, in predict
stonetype_result = stonetype_predict(resized_image)
File "/usr/src/app/services/stonetype_service/app/server.py", line 35, in predict
output = sessionStoneType.run(None, {input_name: image_tensor})
File "/usr/local/lib/python3.10/site-packages/onnxruntime/capi/onnxruntime_inference_collection.py", line 220, in run
return self._sess.run(output_names, input_feed, run_options)
onnxruntime.capi.onnxruntime_pybind11_state.InvalidArgument: [ONNXRuntimeError] : 2 : INVALID_ARGUMENT : Got invalid dimensions for input: images for the following indices
index: 2 Got: 224 Expected: 640
index: 3 Got: 224 Expected: 640
Please fix either the inputs or the model.

解决方案

先排查核心问题:模型会话复用错误

从错误信息看,分类器模型接收到224x224的输入却要求640x640,说明分类器的会话实例被错误复用了YOLO模型的会话,或是Docker中模型加载路径混乱,导致分类器加载的实际是YOLO模型。

先做2项验证:

  1. 进入Docker容器,手动检查stonetype.onnx等分类器模型的路径和正确性,确认模型输入尺寸为224x224
  2. 在load_model函数中添加日志,打印加载模型的输入尺寸:
def load_model(onnx_file_path):
    session = ort.InferenceSession(onnx_file_path)
    input_shape = session.get_inputs()[0].shape
    print(f"Loaded model {onnx_file_path}, input shape: {input_shape}")
    return session

启动容器后观察日志,确认每个模型的输入尺寸符合预期。

多线程实现方案(解决并发与推理效率问题)

若模型加载无问题,可通过多线程优化推理并发能力:

1. 启用Flask多线程模式

修改server.py的启动逻辑,开启多线程处理请求:

if __name__ == "__main__":
    app.run(host="0.0.0.0", port=5000, threaded=True)

或直接修改Docker启动命令:

CMD ["python", "-c", "from server import app; app.run(host='0.0.0.0', port=5000, threaded=True)"]

2. 全局预加载模型会话

ONNX Runtime的InferenceSession线程安全,建议Flask启动时全局加载所有模型,避免每次请求重复加载:

# 全局加载所有模型
yolo_session = load_model("services/yolo_service/yolo.onnx")
stonetype_session = load_model("services/stonetype_service/stonetype.onnx")
cut_session = load_model("services/cut_service/cut.onnx")
color_session = load_model("services/color_service/color.onnx")

# 推理函数使用全局会话
def yolo_predict(image):
    tensor = preprocess_yolo(image)  # 预处理为640x640
    return infer(yolo_session, tensor)

def stonetype_predict(image):
    tensor = preprocess_classifier(image)  # 预处理为224x224
    return infer(stonetype_session, tensor)

3. 并行推理多个分类器

用ThreadPoolExecutor并行处理多个分类器,提升推理速度:

from concurrent.futures import ThreadPoolExecutor

# 全局创建线程池(根据CPU核心数设置max_workers)
executor = ThreadPoolExecutor(max_workers=3)

@app.route("/predict", methods=["POST"])
def predict():
    # ... YOLO推理与图像裁剪 ...
    for box in boxes:
        # 并行提交分类器推理任务
        futures = [
            executor.submit(stonetype_predict, resized_image),
            executor.submit(cut_predict, resized_image),
            executor.submit(color_predict, resized_image)
        ]
        # 获取并行执行结果
        stonetype_result = futures[0].result()
        cut_result = futures[1].result()
        color_result = futures[2].result()
        # ... 结果处理 ...
    return jsonify(results)

Docker资源优化

若存在资源瓶颈,启动容器时指定CPU和内存配额:

docker run -d --name flask-ml-api -p 5000:5000 --cpus=4 --memory=8g your-image-name

同时优化ONNX Runtime的线程配置:

def load_model(onnx_file_path):
    session_options = ort.SessionOptions()
    session_options.intra_op_num_threads = 2  # 算子内并行线程数
    session_options.inter_op_num_threads = 2  # 算子间并行线程数
    session = ort.InferenceSession(onnx_file_path, sess_options=session_options, providers=["CPUExecutionProvider"])
    return session

内容的提问来源于stack exchange,提问作者wadie el

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 09:15:28