在Google Colab加载PEFT LLM模型时内存耗尽如何解决?
问题:加载OpenLLaMA 3B模型时触发内存耗尽错误
运行以下代码时遇到all available RAM has been used错误(未执行微调操作),尝试过关闭8bit量化、移除device_map参数,切换标准GPU和T4 GPU后仍崩溃。是否必须购买Colab Pro?朋友曾用标准Colab完成PEFT操作,求解决方案。
原代码:
model_path = 'openlm-research/open_llama_3b_v2' tokenizer = LlamaTokenizer.from_pretrained(model_path) model = LlamaForCausalLM.from_pretrained(model_path, load_in_8bit=True, device_map = 'auto', ) # Pass in a prompt and infer with the model prompt = 'Q: Create a detailed description for the following product: Corelogic Smooth Mouse, belonging to category: Optical Mouse\nA:' input_ids = tokenizer(prompt, return_tensors="pt").input_ids generation_output = model.generate( input_ids=input_ids, max_new_tokens=128 ) print(tokenizer.decode(generation_output[0]))
解决方案
启用4bit量化(最有效)
4bit量化比8bit更节省内存,且基本不影响推理效果。确保bitsandbytes库版本足够新,加载模型时添加以下配置:import torch from transformers import LlamaTokenizer, LlamaForCausalLM model_path = 'openlm-research/open_llama_3b_v2' tokenizer = LlamaTokenizer.from_pretrained(model_path) model = LlamaForCausalLM.from_pretrained( model_path, load_in_4bit=True, device_map='cuda:0', # 强制加载到GPU,避免CPU内存占用 bnb_4bit_use_double_quant=True, bnb_4bit_quant_type='nf4', bnb_4bit_compute_dtype=torch.bfloat16 )强制模型加载到GPU
不要使用device_map='auto',直接指定device_map='cuda:0',确保模型全部加载到GPU显存,避免拆分到CPU内存导致RAM耗尽。清理Colab环境冗余占用
- 重启Colab会话(菜单栏「Runtime」→「Restart runtime」),清除内存残留。
- 关闭浏览器闲置标签页,减少系统内存占用。
- 运行
!free -h查看内存使用情况,排查其他进程占用。
优化推理阶段内存
生成文本时添加pad_token_id(LlamaTokenizer默认无pad token,用eos token替代),确保输入张量在GPU上:generation_output = model.generate( input_ids=input_ids.to('cuda'), max_new_tokens=128, pad_token_id=tokenizer.eos_token_id, use_cache=True # 保持开启提升速度,内存仍不足可设为False(会变慢) )更新依赖版本
确保transformers >= 4.28.0、bitsandbytes >= 0.39.0,旧版本可能存在内存管理bug,运行以下命令更新:!pip install --upgrade transformers bitsandbytes accelerate
以上方法在标准Colab的T4 GPU(16GB显存)上即可运行OpenLLaMA 3B推理,无需购买Colab Pro。
内容的提问来源于stack exchange,提问作者Obi Anthony
相关产品推荐
相关产品推荐

