如何将已实例化的Hugging Face模型加载到vLLM中进行推理?
问题描述
我正在做一个项目,需要将已经通过Hugging Face Transformers加载并部署在GPU上的模型直接传入vLLM框架做推理,不想从磁盘或模型中心重新加载。
目前用Hugging Face加载模型的代码:
from transformers import GPT2LMHeadModel # 加载已微调且存在于内存中的GPT-2模型 model = GPT2LMHeadModel.from_pretrained('gpt2') model.to('cuda') # 模型已移至GPU
我尝试直接把实例化的模型传给vLLM的LLM类,但代码无法运行:
from transformers import GPT2LMHeadModel, GPT2Tokenizer # 步骤1:使用Hugging Face加载并保存模型 model_name = 'gpt2' model = GPT2LMHeadModel.from_pretrained(model_name) tokenizer = GPT2Tokenizer.from_pretrained(model_name) # 保存模型与分词器 model_save_path = './model/gpt2' tokenizer.save_pretrained(model_save_path) model.save_pretrained(model_save_path) # 步骤2:使用保存的模型路径在vLLM中加载模型 from vllm import LLM, SamplingParams # 尝试直接传入已实例化模型 vllm_model = LLM(model=model) # 定义采样参数与提示词 sampling_params = SamplingParams(temperature=0.8, top_p=0.95) prompts = ["Hello, my name is", "The future of AI is"] # 在vLLM中使用指定模型与参数生成文本 outputs = vllm_model.generate(prompts, sampling_params) # 打印生成的文本 for output in outputs: print(f"Prompt: {output.prompt}, Generated text: {output.outputs[0].text}")
报错信息:
Exception has occurred: OSError Incorrect path_or_model_id: 'GPT2LMHeadModel( (transformer): GPT2Model( (wte): Embedding(50257, 768) (wpe): Embedding(1024, 768) (drop): Dropout(p=0.1, inplace=False) (h): ModuleList( (0-11): 12 x GPT2Block( (ln_1): LayerNorm((768,), eps=1e-05, elementwise_affine=True) (attn): GPT2Attention( (c_attn): Conv1D() (c_proj): Conv1D() (attn_dropout): Dropout(p=0.1, inplace=False) (resid_dropout): Dropout(p=0.1, inplace=False) ) (ln_2): LayerNorm((768,), eps=1e-05, elementwise_affine=True) (mlp): GPT2MLP( (c_fc): Conv1D() (c_proj): Conv1D() (act): NewGELUActivation() (dropout): Dropout(p=0.1, inplace=False) ) ) ) (ln_f): LayerNorm((768,), eps=1e-05, elementwise_affine=True) ) (lm_head): Linear(in_features=768, out_features=50257, bias=False) )'. Please provide either the path to a local folder or the repo_id of a model on the Hub. huggingface_hub.errors.HFValidationError: Repo id must use alphanumeric chars or '-', '_', '.', '--' and '..' are forbidden, '-' and '.' cannot start or end the name, max length is 96: 'GPT2LMHeadModel( (transformer): GPT2Model( (wte): Embedding(50257, 768) (wpe): Embedding(1024, 768) (drop): Dropout(p=0.1, inplace=False) (h): ModuleList( (0-11): 12 x GPT2Block( (ln_1): LayerNorm((768,), eps=1e-05, elementwise_affine=True) (attn): GPT2Attention( (c_attn): Conv1D() (c_proj): Conv1D() (attn_dropout): Dropout(p=0.1, inplace=False) (resid_dropout): Dropout(p=0.1, inplace=False) ) (ln_2): LayerNorm((768,), eps=1e-05, elementwise_affine=True) (mlp): GPT2MLP( (c_fc): Conv1D() (c_proj): Conv1D() (act): NewGELUActivation() (dropout): Dropout(p=0.1, inplace=False) ) ) ) (ln_f): LayerNorm((768,), eps=1e-05, elementwise_affine=True) ) (lm_head): Linear(in_features=768, out_features=50257, bias=False) )'. The above exception was the direct cause of the following exception: File "/lfs/ampere1/0/brando9/gold-ai-olympiad/py_src/training/maf_self_improv_train.py", line 35, in <module> vllm_model = LLM(model=model) ^^^^^^^^^^^^^^^^ OSError: Incorrect path_or_model_id: 'GPT2LMHeadModel( (transformer): GPT2Model( (wte): Embedding(50257, 768) (wpe): Embedding(1024, 768) (drop): Dropout(p=0.1, inplace=False) (h): ModuleList( (0-11): 12 x GPT2Block( (ln_1): LayerNorm((768,), eps=1e-05, elementwise_affine=True) (attn): GPT2Attention( (c_attn): Conv1D() (c_proj): Conv1D() (attn_dropout): Dropout(p=0.1, inplace=False) (resid_dropout): Dropout(p=0.1, inplace=False) ) (ln_2): LayerNorm((768,), eps=1e-05, elementwise_affine=True) (mlp): GPT2MLP( (c_fc): Conv1D() (c_proj): Conv1D() (act): NewGELUActivation() (dropout): Dropout(p=0.1, inplace=False) ) ) ) (ln_f): LayerNorm((768,), eps=1e-05, elementwise_affine=True) ) (lm_head): Linear(in_features=768, out_features=50257, bias=False) )'. Please provide either the path to a local folder or the repo_id of a model on the Hub.
解决方案
核心原因
vLLM的LLM类设计上仅接受模型路径或Hugging Face Hub的模型ID作为model参数,不支持直接传入已实例化的Hugging Face模型对象。这是因为vLLM需要自行管理模型内存、应用PagedAttention等优化,必须重新加载模型到其专属的内存空间。
可行替代方案
1. 临时保存到内存文件系统(最快实现)
将已加载的模型临时保存到/tmp这类内存文件系统,再让vLLM从该路径加载,避免磁盘IO开销:
import tempfile from transformers import GPT2LMHeadModel, GPT2Tokenizer from vllm import LLM, SamplingParams # 加载已有的HF模型 model = GPT2LMHeadModel.from_pretrained('gpt2') model.to('cuda') tokenizer = GPT2Tokenizer.from_pretrained('gpt2') # 创建临时目录保存模型 with tempfile.TemporaryDirectory() as tmp_dir: model.save_pretrained(tmp_dir) tokenizer.save_pretrained(tmp_dir) # vLLM从临时目录加载模型 vllm_model = LLM(model=tmp_dir) # 推理 sampling_params = SamplingParams(temperature=0.8, top_p=0.95) prompts = ["Hello, my name is", "The future of AI is"] outputs = vllm_model.generate(prompts, sampling_params) for output in outputs: print(f"Prompt: {output.prompt}, Generated text: {output.outputs[0].text}")
临时目录会在代码块结束后自动删除,不会残留文件。
2. 自定义模型加载(进阶)
如果不想保存模型,可以修改vLLM的底层加载逻辑,直接导入模型权重。但这种方法需要熟悉vLLM的源码结构,且兼容性较差,不推荐用于生产环境。具体可以参考vLLM的HuggingFaceModel类实现,手动构建vLLM的模型实例并加载已有的权重。
注意事项
- vLLM对模型结构有兼容性要求,确保你的HF模型是vLLM支持的架构(比如GPT2、LLaMA等)
- 临时保存方案虽然需要一次序列化/反序列化,但内存文件系统的速度足够快,不会成为性能瓶颈
- 如果模型是微调后的,保存时要确保所有权重都被正确保存(包括LoRA权重,如果使用了LoRA)
内容的提问来源于stack exchange,提问作者Charlie Parker
相关产品推荐
相关产品推荐

