使用QLoRA微调LLaMA 3.1 8B时遇CUDA内存不足错误求助
问题描述
使用4-bit bitsandbytes库结合QLoRA技术微调Hugging Face的LLaMA 3.1 8B-Instruct模型,训练数据集为Amod/mental_health_counseling_conversations,运行代码时触发torch.cuda.OutOfMemoryError。已尝试使用多GPU、提升GPU显存、调整batch size,问题仍未解决。
代码片段
import torch from torch.utils.data import Dataset, DataLoader from transformers import Trainer, AutoModelForCausalLM, AutoTokenizer from datasets import Dataset, load_dataset from peft import get_peft_model, LoraConfig, TaskType import numpy as np from transformers import BitsAndBytesConfig, TrainingArguments # BitsAndBytes configuration which loads in 4-bit bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_quant_storage=torch.bfloat16, ) # Load model and tokenizer using huggingface model = AutoModelForCausalLM.from_pretrained( "meta-llama/Meta-Llama-3.1-8B-Instruct", quantization_config=bnb_config, torch_dtype=torch.bfloat16, ) tokenizer = AutoTokenizer.from_pretrained("meta-llama/Meta-Llama-3.1-8B-Instruct") # Load mental health counseling dataset dataset = load_dataset("Amod/mental_health_counseling_conversations") # Data preprocessing functions def generate_prompt(Context, Response): return f""" You are supposed to reply to the questions as a professional therapist Question: {Context} Answer: {Response} """ def format_for_llama(example): prompt = generate_prompt(example['Context'], example['Response']) return { "text": prompt.strip() } formatted_dataset = dataset['train'].map(format_for_llama) tokenizer.pad_token = tokenizer.eos_token # Collate function for DataLoader def collate_fn(examples): input_ids = torch.stack([example['input_ids'] for example in examples]) attention_mask = torch.stack([example['attention_mask'] for example in examples]) return { 'input_ids': input_ids, 'attention_mask': attention_mask } train_dataloader = DataLoader(tokenized_dataset, collate_fn=collate_fn, batch_size=10) # PEFT configuration (adding trainable adapters) peft_config = LoraConfig( task_type=TaskType.CAUSAL_LM, inference_mode=False, r=16, lora_alpha=32, lora_dropout=0.1, target_modules=[ "q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj" ] ) model = get_peft_model(model, peft_config) model.print_trainable_parameters() # Training arguments hyperparameters args = TrainingArguments( output_dir="./models", save_strategy="epoch", learning_rate=2e-5, per_device_train_batch_size=10, num_train_epochs=10, weight_decay=0.01, logging_dir='logs', logging_strategy="epoch", remove_unused_columns=False, eval_strategy="no", load_best_model_at_end=False, ) # Trainer initialization and training trainer = Trainer( model=model, args=args, train_dataset=tokenized_dataset, data_collator=collate_fn ) trainer.train()
错误日志
OutOfMemoryError Traceback (most recent call last) Cell In[42], line 7 1 trainer = Trainer( 2 model=model, 3 args=args, 4 train_dataset=tokenized_dataset, 5 data_collator=collate_fn 6 ) ----> 7 trainer.train() File /usr/local/lib/python3.10/dist-packages/transformers/trainer.py:1938, in Trainer.train(self, resume_from_checkpoint, trial, ignore_keys_for_eval, **kwargs) 1936 hf_hub_utils.enable_progress_bars() 1937 else: -> 1938 return inner_training_loop( 1939 args=args, 1940 resume_from_checkpoint=resume_from_checkpoint, 1941 trial=trial, 1942 ignore_keys_for_eval=ignore_keys_for_eval, 1943 ) File /usr/local/lib/python3.10/dist-packages/transformers/trainer.py:2279, in Trainer._inner_training_loop(self, batch_size, args, resume_from_checkpoint, trial, ignore_keys_for_eval) 2276 self.control = self.callback_handler.on_step_begin(args, self.state, self.control) 2278 with self.accelerator.accumulate(model): -> 2279 tr_loss_step = self.training_step(model, inputs) ... 1857 else: -> 1858 ret = input.softmax(dim, dtype=dtype) 1859 return ret OutOfMemoryError: CUDA out of memory. Tried to allocate 1.25 GiB. GPU 0 has a total capacty of 47.54 GiB of which 1.05 GiB is free. Process 3704361 has 46.47 GiB memory in use. Of the allocated memory 45.37 GiB is allocated by PyTorch, and 808.03 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting max_split_size_mb to avoid fragmentation. See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF
已尝试方案
- 使用多GPU
- 增加GPU显存
- 调整batch size
环境信息
- Python版本:3.10
- PyTorch版本:待指定
- CUDA版本:待指定
- GPU:RunPod多实例
解决方案
1. 修复数据预处理缺失环节
代码中定义了formatted_dataset但未执行tokenization就直接使用tokenized_dataset,属于逻辑错误。添加tokenization步骤并限制序列长度,减少单样本显存占用:
def tokenize_function(examples): return tokenizer( examples["text"], truncation=True, max_length=512, # 限制最大序列长度,可根据数据集调整 padding="max_length" ) # 执行tokenization并转换为PyTorch张量格式 tokenized_dataset = formatted_dataset.map(tokenize_function, batched=True) tokenized_dataset.set_format("torch", columns=["input_ids", "attention_mask"]) # 删除手动创建的train_dataloader,Trainer会自动处理数据加载
2. 优化训练参数配置
调整TrainingArguments,用梯度累积弥补小batch size的不足,同时启用显存优化选项:
args = TrainingArguments( output_dir="./models", save_strategy="steps", save_steps=1000, # 降低模型保存频率,减少显存瞬时占用 learning_rate=2e-5, per_device_train_batch_size=2, # 进一步降低单卡batch size gradient_accumulation_steps=8, # 梯度累积维持有效batch size(2*8=16,接近原batch size) num_train_epochs=10, weight_decay=0.01, logging_dir='logs', logging_strategy="epoch", remove_unused_columns=False, eval_strategy="no", load_best_model_at_end=False, gradient_checkpointing=True, # 启用梯度 checkpointing,用计算量换显存 bf16=True, # 启用bf16混合精度(GPU支持时),否则改为fp16=True fp16=False, )
3. 强化BitsAndBytes量化效果
启用双重量化进一步压缩模型显存占用:
bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_quant_storage=torch.bfloat16, bnb_4bit_use_double_quant=True, # 启用双重量化 )
4. 优化多GPU训练流程
使用accelerate工具启动训练,确保模型正确分配到多GPU:
- 执行
accelerate config生成多GPU配置文件 - 用
accelerate launch your_script.py启动训练脚本
5. 清理显存碎片
在训练前添加显存清理代码,并设置显存分配策略:
import os os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "max_split_size_mb:128" # 减少显存碎片 import gc torch.cuda.empty_cache() gc.collect()
6. 减少LoRA目标模块(可选)
如果显存仍紧张,可减少LoRA训练的目标模块,降低可训练参数数量:
peft_config = LoraConfig( task_type=TaskType.CAUSAL_LM, inference_mode=False, r=16, lora_alpha=32, lora_dropout=0.1, target_modules=[ "q_proj", "v_proj" # 仅训练注意力层的关键模块 ] )
内容的提问来源于stack exchange,提问作者Palthya Deepmalik
相关产品推荐
相关产品推荐

