You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将已实例化的Hugging Face模型加载到vLLM中进行推理?

问题描述

我正在做一个项目,需要将已经通过Hugging Face Transformers加载并部署在GPU上的模型直接传入vLLM框架做推理,不想从磁盘或模型中心重新加载。

目前用Hugging Face加载模型的代码:

from transformers import GPT2LMHeadModel

# 加载已微调且存在于内存中的GPT-2模型
model = GPT2LMHeadModel.from_pretrained('gpt2')
model.to('cuda')  # 模型已移至GPU

我尝试直接把实例化的模型传给vLLM的LLM类,但代码无法运行:

from transformers import GPT2LMHeadModel, GPT2Tokenizer

# 步骤1:使用Hugging Face加载并保存模型
model_name = 'gpt2'
model = GPT2LMHeadModel.from_pretrained(model_name)
tokenizer = GPT2Tokenizer.from_pretrained(model_name)

# 保存模型与分词器
model_save_path = './model/gpt2'
tokenizer.save_pretrained(model_save_path)
model.save_pretrained(model_save_path)

# 步骤2:使用保存的模型路径在vLLM中加载模型
from vllm import LLM, SamplingParams

# 尝试直接传入已实例化模型
vllm_model = LLM(model=model)

# 定义采样参数与提示词
sampling_params = SamplingParams(temperature=0.8, top_p=0.95)
prompts = ["Hello, my name is", "The future of AI is"]

# 在vLLM中使用指定模型与参数生成文本
outputs = vllm_model.generate(prompts, sampling_params)

# 打印生成的文本
for output in outputs:
    print(f"Prompt: {output.prompt}, Generated text: {output.outputs[0].text}")

报错信息:

Exception has occurred: OSError
Incorrect path_or_model_id: 'GPT2LMHeadModel(
  (transformer): GPT2Model(
    (wte): Embedding(50257, 768)
    (wpe): Embedding(1024, 768)
    (drop): Dropout(p=0.1, inplace=False)
    (h): ModuleList(
      (0-11): 12 x GPT2Block(
        (ln_1): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
        (attn): GPT2Attention(
          (c_attn): Conv1D()
          (c_proj): Conv1D()
          (attn_dropout): Dropout(p=0.1, inplace=False)
          (resid_dropout): Dropout(p=0.1, inplace=False)
        )
        (ln_2): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
        (mlp): GPT2MLP(
          (c_fc): Conv1D()
          (c_proj): Conv1D()
          (act): NewGELUActivation()
          (dropout): Dropout(p=0.1, inplace=False)
        )
      )
    )
    (ln_f): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
  )
  (lm_head): Linear(in_features=768, out_features=50257, bias=False)
)'. Please provide either the path to a local folder or the repo_id of a model on the Hub.
huggingface_hub.errors.HFValidationError: Repo id must use alphanumeric chars or '-', '_', '.', '--' and '..' are forbidden, '-' and '.' cannot start or end the name, max length is 96: 'GPT2LMHeadModel(
  (transformer): GPT2Model(
    (wte): Embedding(50257, 768)
    (wpe): Embedding(1024, 768)
    (drop): Dropout(p=0.1, inplace=False)
    (h): ModuleList(
      (0-11): 12 x GPT2Block(
        (ln_1): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
        (attn): GPT2Attention(
          (c_attn): Conv1D()
          (c_proj): Conv1D()
          (attn_dropout): Dropout(p=0.1, inplace=False)
          (resid_dropout): Dropout(p=0.1, inplace=False)
        )
        (ln_2): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
        (mlp): GPT2MLP(
          (c_fc): Conv1D()
          (c_proj): Conv1D()
          (act): NewGELUActivation()
          (dropout): Dropout(p=0.1, inplace=False)
        )
      )
    )
    (ln_f): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
  )
  (lm_head): Linear(in_features=768, out_features=50257, bias=False)
)'.

The above exception was the direct cause of the following exception:

  File "/lfs/ampere1/0/brando9/gold-ai-olympiad/py_src/training/maf_self_improv_train.py", line 35, in <module>
    vllm_model = LLM(model=model)
                 ^^^^^^^^^^^^^^^^
OSError: Incorrect path_or_model_id: 'GPT2LMHeadModel(
  (transformer): GPT2Model(
    (wte): Embedding(50257, 768)
    (wpe): Embedding(1024, 768)
    (drop): Dropout(p=0.1, inplace=False)
    (h): ModuleList(
      (0-11): 12 x GPT2Block(
        (ln_1): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
        (attn): GPT2Attention(
          (c_attn): Conv1D()
          (c_proj): Conv1D()
          (attn_dropout): Dropout(p=0.1, inplace=False)
          (resid_dropout): Dropout(p=0.1, inplace=False)
        )
        (ln_2): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
        (mlp): GPT2MLP(
          (c_fc): Conv1D()
          (c_proj): Conv1D()
          (act): NewGELUActivation()
          (dropout): Dropout(p=0.1, inplace=False)
        )
      )
    )
    (ln_f): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
  )
  (lm_head): Linear(in_features=768, out_features=50257, bias=False)
)'. Please provide either the path to a local folder or the repo_id of a model on the Hub.
解决方案

核心原因

vLLM的LLM类设计上仅接受模型路径或Hugging Face Hub的模型ID作为model参数,不支持直接传入已实例化的Hugging Face模型对象。这是因为vLLM需要自行管理模型内存、应用PagedAttention等优化,必须重新加载模型到其专属的内存空间。

可行替代方案

1. 临时保存到内存文件系统(最快实现)

将已加载的模型临时保存到/tmp这类内存文件系统,再让vLLM从该路径加载,避免磁盘IO开销:

import tempfile
from transformers import GPT2LMHeadModel, GPT2Tokenizer
from vllm import LLM, SamplingParams

# 加载已有的HF模型
model = GPT2LMHeadModel.from_pretrained('gpt2')
model.to('cuda')
tokenizer = GPT2Tokenizer.from_pretrained('gpt2')

# 创建临时目录保存模型
with tempfile.TemporaryDirectory() as tmp_dir:
    model.save_pretrained(tmp_dir)
    tokenizer.save_pretrained(tmp_dir)
    
    # vLLM从临时目录加载模型
    vllm_model = LLM(model=tmp_dir)
    
    # 推理
    sampling_params = SamplingParams(temperature=0.8, top_p=0.95)
    prompts = ["Hello, my name is", "The future of AI is"]
    outputs = vllm_model.generate(prompts, sampling_params)
    
    for output in outputs:
        print(f"Prompt: {output.prompt}, Generated text: {output.outputs[0].text}")

临时目录会在代码块结束后自动删除,不会残留文件。

2. 自定义模型加载(进阶)

如果不想保存模型,可以修改vLLM的底层加载逻辑,直接导入模型权重。但这种方法需要熟悉vLLM的源码结构,且兼容性较差,不推荐用于生产环境。具体可以参考vLLM的HuggingFaceModel类实现,手动构建vLLM的模型实例并加载已有的权重。

注意事项

  • vLLM对模型结构有兼容性要求,确保你的HF模型是vLLM支持的架构(比如GPT2、LLaMA等)
  • 临时保存方案虽然需要一次序列化/反序列化,但内存文件系统的速度足够快,不会成为性能瓶颈
  • 如果模型是微调后的,保存时要确保所有权重都被正确保存(包括LoRA权重,如果使用了LoRA)

内容的提问来源于stack exchange,提问作者Charlie Parker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 00:44:54