You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PeftModel加载Hugging Face的Gemma-7B微调模型时遇InvalidHeaderDeserialization错误

问题:加载Hugging Face上的Gemma-7B LoRA微调模型报错

我基于一份Jupyter Notebook微调了google/gemma-7b模型,仅修改了一处:训练结束后通过trainer.push_to_hub()将模型保存至Hugging Face。但尝试用以下代码从Hugging Face加载我的模型时:

import torch
from peft import LoraConfig, PeftModel
from transformers import AutoModelForCausalLM, BitsAndBytesConfig, AutoTokenizer

base_model = "google/gemma-7b"
my_model = "rreit/gemma-7b-prompts"

model = AutoModelForCausalLM.from_pretrained(
    base_model,
    load_in_8bit=True,
    torch_dtype=torch.float16,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(base_model)

model = PeftModel.from_pretrained(model, my_model)

def generate_and_tokenize_prompt(data_point):
    eval_prompt =f"""You are a powerful text-to-C# model. Your job is to answer questions about a FONS Enterprise.You are given a question and context of the file.
    
    You must output the code that suits the question.

    ### Input:
    {data_point["prompt"]}

    ### Context:
    {data_point["context"]}

    ### Response:
    """
    model_input = tokenizer(eval_prompt, return_tensors="pt").to("cuda")

    model.eval()
    with torch.no_grad():
        print(tokenizer.decode(model.generate(**model_input, max_new_tokens=512)[0], skip_special_tokens=True))

with open("context.cs","r") as f:
    context = f.read()

data = {
    "prompt": 'Write a method to extend the validation of a PharmacyBO object, ensuring that the "Code" attribute excludes the character',
    "context": context
}

generate_and_tokenize_prompt(data)

出现错误:safetensors_rust.SafetensorError: Error while deserializing header: InvalidHeaderDeserialization

错误堆栈信息如下:

The `load_in_4bit` and `load_in_8bit` arguments are deprecated and will be removed in the future versions. Please, pass a `BitsAndBytesConfig` object in `quantization_config` argument instead.
Loading checkpoint shards: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:05<00:00,  1.30s/it]
Traceback (most recent call last):
  File "run_prompt.py", line 21, in <module>
    model = PeftModel.from_pretrained(model, 'prompt-gemma-7B/')
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "peft_model.py", line 353, in from_pretrained
    model.load_adapter(model_id, adapter_name, is_trainable=is_trainable, **kwargs)
  File "peft_model.py", line 694, in load_adapter
    adapters_weights = load_peft_weights(model_id, device=torch_device, **hf_hub_download_kwargs)
                       ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "utils/save_and_load.py", line 326, in load_peft_weights
    adapters_weights = safe_load_file(filename, device=device)
                       ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "pip_packages/lib/python3.11/site-packages/safetensors/torch.py", line 308, in load_file
    with safe_open(filename, framework="pt", device=device) as f:
         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
safetensors_rust.SafetensorError: Error while deserializing header: InvalidHeaderDeserialization

我可以通过torch.load('prompt-gemma-7B/checkpoint-2000/optimizer.pt')从本地checkpoint正常加载模型,代码如下:

import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig, AutoTokenizer, Trainer, AutoModel
from peft import (
    LoraConfig,
    get_peft_model,
    get_peft_model_state_dict,
    prepare_model_for_int8_training,
    set_peft_model_state_dict,
)

base_model = "google/gemma-7b"
my_model = "rreit/gemma-7b-prompts"

model = AutoModelForCausalLM.from_pretrained(
    base_model,
    load_in_8bit=True,
    torch_dtype=torch.float16,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(base_model)

model = prepare_model_for_int8_training(model)

config = LoraConfig(
    r=16,
    lora_alpha=16,
    target_modules=[
    "q_proj",
    "k_proj",
    "v_proj",
    "o_proj",
],
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM",
)
model = get_peft_model(model, config)

adapters_weights = torch.load('prompt-gemma-7B/checkpoint-2000/optimizer.pt')
set_peft_model_state_dict(model, adapters_weights)

但无法通过PeftModel.from_pretrained()从Hugging Face加载该模型。另外,移除以下代码重新训练后模型可正常加载:

model.config.use_cache = False

old_state_dict = model.state_dict
model.state_dict = (lambda self, *_, **__: get_peft_model_state_dict(self, old_state_dict())).__get__(
    model, type(model)
)
if torch.__version__ >= "2" and sys.platform != "win32":
    print("compiling the model")
    model = torch.compile(model)

我也尝试用以下代码重新上传模型至Hugging Face,但问题依旧:

import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig, AutoTokenizer, Trainer, AutoModel
from peft import (
    LoraConfig,
    get_peft_model,
    get_peft_model_state_dict,
    prepare_model_for_int8_training,
    set_peft_model_state_dict,
)

base_model = "google/gemma-7b"
model = AutoModelForCausalLM.from_pretrained(
    base_model,
    load_in_8bit=True,
    torch_dtype=torch.float16,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("google/gemma-7b")

model = prepare_model_for_int8_training(model)

config = LoraConfig(
    r=16,
    lora_alpha=16,
    target_modules=[
    "q_proj",
    "k_proj",
    "v_proj",
    "o_proj",
],
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM",
)
model = get_peft_model(model, config)

adapters_weights = torch.load('prompt-gemma-7B/checkpoint-2000/optimizer.pt')
set_peft_model_state_dict(model, adapters_weights)

trainer = Trainer(
    model=model,
    )

trainer.push_to_hub("rreit/gemma-7b-prompts")

请问还有其他方法可以从Hugging Face加载该模型吗?


内容的提问来源于stack exchange,提问作者Rreit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 11:17:06