You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Docker+vLLM部署Hugging Face模型遇配置及未知模型错误求助

问题:vLLM v0.8.2部署Qwen2.5-Coder-32B时模型识别失败

使用Docker搭配vLLM v0.8.2部署Hugging Face模型,测试Mistral 7B v0.2、v0.3正常,但部署Qwen2.5-Coder-32B时触发模型识别错误。

所用Dockerfile配置

FROM vllm/vllm-openai:v0.8.2

#Override the entrypoint to run the vLLM server with your model
ENTRYPOINT ["python3", "-m", "vllm.entrypoints.openai.api_server", \
  "--model", "Qwen/Qwen2.5-Coder-32B", \
  "--host", "0.0.0.0", \
  "--port", "7860", \
  "--tensor-parallel-size", "4", \
  "--trust-remote-code"]

错误日志

===== Application Startup at 2025-03-31 20:38:05 =====
INFO 03-31 13:39:57 [__init__.py:239] Automatically detected platform cuda.
INFO 03-31 13:39:59 [api_server.py:981] vLLM API server version 0.8.2
...
ValueError: Unrecognized model in Qwen/Qwen2.5-Coder-32B. Should have a `model_type` key in its config.json, or contain one of the following strings in its name: ...

解决方案

  • 核心原因:vLLM v0.8.2发布早于Qwen2.5系列模型,原生未适配该模型的model_type,且模型名称中没有v0.8.2能识别的关键词。

  • 方案1:升级vLLM版本(推荐)
    将Dockerfile中的基础镜像替换为支持Qwen2.5的vLLM版本(如v0.9.0及以上),修改后Dockerfile的第一行为:

    FROM vllm/vllm-openai:v0.9.0
    

    升级后无需额外调整启动参数,vLLM会自动识别Qwen2.5模型。

  • 方案2:在v0.8.2中强制指定模型类型
    如果必须保留v0.8.2版本,在启动命令中添加--model-type qwen2参数,强制vLLM使用Qwen2的模型处理逻辑,同时确保--trust-remote-code参数已开启。修改后的ENTRYPOINT如下:

    ENTRYPOINT ["python3", "-m", "vllm.entrypoints.openai.api_server", \
      "--model", "Qwen/Qwen2.5-Coder-32B", \
      "--host", "0.0.0.0", \
      "--port", "7860", \
      "--tensor-parallel-size", "4", \
      "--trust-remote-code", \
      "--model-type", "qwen2"]
    

内容的提问来源于stack exchange,提问作者Marios Petrov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 14:59:50