运行Law-GPT项目遇ValueError:需用disk_offload替代全模型磁盘卸载
解决Law-GPT加载Llama-2模型时的磁盘卸载错误
问题重现
运行Law-GPT项目加载meta-llama/Llama-2-7b-chat-hf模型时,设置offload_folder和offload_state_dict参数后触发以下错误:
ValueError: You are trying to offload the whole model to the disk. Please use the `disk_offload` function instead.
原因分析
当device_map='auto'但本地GPU显存完全不足以承载模型的任何部分时,Transformers会尝试将整个模型卸载到磁盘,但此时不能通过offload_folder参数配置,必须使用专门的disk_offload方法。
解决方案
方案1:开启4bit量化(优先推荐)
4bit量化能大幅降低模型显存占用,让模型可以在普通GPU上运行,无需全量磁盘卸载。修改load_llm函数如下:
def load_llm(): """ Load the LLM """ # Model ID repo_id = 'meta-llama/Llama-2-7b-chat-hf' login(token="hf_xxxxxxxx") # Load the model with 4bit quantization model = AutoModelForCausalLM.from_pretrained( repo_id, device_map='auto', load_in_4bit=True, # 开启4bit量化 bnb_4bit_compute_dtype=torch.float16, # 搭配float16计算加速 token = True ) # Load the tokenizer tokenizer = AutoTokenizer.from_pretrained( repo_id, use_fast=True ) # Create pipeline pipe = pipeline( 'text-generation', model=model, tokenizer=tokenizer, max_length=512 ) # Load the LLM llm = HuggingFacePipeline(pipeline=pipe) return llm
注意:需要提前安装bitsandbytes库:pip install bitsandbytes
方案2:使用disk_offload全量磁盘加载
如果确实没有足够GPU显存,需要全量卸载到磁盘,修改模型加载逻辑如下:
from accelerate import disk_offload def load_llm(): """ Load the LLM with full disk offload """ # Model ID repo_id = 'meta-llama/Llama-2-7b-chat-hf' login(token="hf_xxxxxxxx") # 先加载模型到CPU model = AutoModelForCausalLM.from_pretrained( repo_id, device_map='cpu', token = True ) # 执行磁盘卸载 disk_offload(model, offload_dir=r"C:\Users\DHRUV\Desktop\New folder\Law-GPT") # Load the tokenizer tokenizer = AutoTokenizer.from_pretrained( repo_id, use_fast=True ) # Create pipeline pipe = pipeline( 'text-generation', model=model, tokenizer=tokenizer, max_length=512 ) # Load the LLM llm = HuggingFacePipeline(pipeline=pipe) return llm
这种方式运行速度会很慢,仅适合无GPU的应急场景。
方案3:手动指定device_map分层加载
如果GPU能承载部分模型层,可以手动指定哪些层放GPU,哪些放CPU/磁盘:
def load_llm(): """ Load the LLM with manual device mapping """ # Model ID repo_id = 'meta-llama/Llama-2-7b-chat-hf' login(token="hf_xxxxxxxx") # 手动指定设备映射,例如前20层放GPU,其余放CPU device_map = { "model.embed_tokens": 0, "model.layers.0": 0, "model.layers.1": 0, # ... 按需添加更多GPU层 "model.layers.19": 0, "model.layers.20": "cpu", # ... 按需添加更多CPU层 "model.norm": "cpu", "lm_head": "cpu" } # 加载模型 model = AutoModelForCausalLM.from_pretrained( repo_id, device_map=device_map, offload_folder=r"C:\Users\DHRUV\Desktop\New folder\Law-GPT", token = True ) # Load the tokenizer tokenizer = AutoTokenizer.from_pretrained( repo_id, use_fast=True ) # Create pipeline pipe = pipeline( 'text-generation', model=model, tokenizer=tokenizer, max_length=512 ) # Load the LLM llm = HuggingFacePipeline(pipeline=pipe) return llm
可以根据自己GPU显存大小调整分层数量,平衡速度和显存占用。
内容的提问来源于stack exchange,提问作者user17589423
相关产品推荐
相关产品推荐

