使用Docker+vLLM部署Hugging Face模型遇配置及未知模型错误求助
问题:vLLM v0.8.2部署Qwen2.5-Coder-32B时模型识别失败
使用Docker搭配vLLM v0.8.2部署Hugging Face模型,测试Mistral 7B v0.2、v0.3正常,但部署Qwen2.5-Coder-32B时触发模型识别错误。
所用Dockerfile配置
FROM vllm/vllm-openai:v0.8.2 #Override the entrypoint to run the vLLM server with your model ENTRYPOINT ["python3", "-m", "vllm.entrypoints.openai.api_server", \ "--model", "Qwen/Qwen2.5-Coder-32B", \ "--host", "0.0.0.0", \ "--port", "7860", \ "--tensor-parallel-size", "4", \ "--trust-remote-code"]
错误日志
===== Application Startup at 2025-03-31 20:38:05 ===== INFO 03-31 13:39:57 [__init__.py:239] Automatically detected platform cuda. INFO 03-31 13:39:59 [api_server.py:981] vLLM API server version 0.8.2 ... ValueError: Unrecognized model in Qwen/Qwen2.5-Coder-32B. Should have a `model_type` key in its config.json, or contain one of the following strings in its name: ...
解决方案
核心原因:vLLM v0.8.2发布早于Qwen2.5系列模型,原生未适配该模型的
model_type,且模型名称中没有v0.8.2能识别的关键词。方案1:升级vLLM版本(推荐)
将Dockerfile中的基础镜像替换为支持Qwen2.5的vLLM版本(如v0.9.0及以上),修改后Dockerfile的第一行为:FROM vllm/vllm-openai:v0.9.0升级后无需额外调整启动参数,vLLM会自动识别Qwen2.5模型。
方案2:在v0.8.2中强制指定模型类型
如果必须保留v0.8.2版本,在启动命令中添加--model-type qwen2参数,强制vLLM使用Qwen2的模型处理逻辑,同时确保--trust-remote-code参数已开启。修改后的ENTRYPOINT如下:ENTRYPOINT ["python3", "-m", "vllm.entrypoints.openai.api_server", \ "--model", "Qwen/Qwen2.5-Coder-32B", \ "--host", "0.0.0.0", \ "--port", "7860", \ "--tensor-parallel-size", "4", \ "--trust-remote-code", \ "--model-type", "qwen2"]
内容的提问来源于stack exchange,提问作者Marios Petrov
相关产品推荐
相关产品推荐

