使用双RTX4090微调Llama-2-7b时CUDA内存不足问题排查求助
问题描述
我使用两块RTX 4090 GPU微调Llama-2-7b-chat-hf模型,多次尝试后仍持续出现"CUDA out of memory"错误。希望确定该问题源于GPU资源限制还是代码存在低效/错误,附上使用的代码,盼定位内存问题根源。考虑过数据并行或模型剪枝,但缺乏相关知识。
代码片段
import os import torch from datasets import load_dataset from transformers import ( AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig, HfArgumentParser, TrainingArguments, pipeline, logging, TextDataset, DataCollatorForLanguageModeling, Trainer ) from datasets import Dataset from peft import LoraConfig, PeftModel from trl import SFTTrainer import pandas as pd import json import gc os.environ["CUDA_VISIBLE_DEVICES"] = "0,1" gc.collect() os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "max_split_size_mb:40000" base_model = r"***" new_model = r"***" json_dataset = load_dataset("json", data_files="***") json_dataset = pd.DataFrame(json_dataset['train']) device = torch.device("cuda") if torch.cuda.is_available() else torch.device("cpu") bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_use_double_quant=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16 ) model = AutoModelForCausalLM.from_pretrained(base_model) tokenizer = AutoTokenizer.from_pretrained(base_model, trust_remote_code=True) tokenizer.pad_token = tokenizer.eos_token tokenizer.padding_side = "right" model.config.pretraining_tp = 1 tokenized_dataset = updated_dataset.map( tokenize_function, batched=True, batch_size=1, drop_last_batch=True ) print(tokenized_dataset) tokenized_dataset = tokenized_dataset.add_column("labels", tokenized_dataset["input_ids"]) print(tokenized_dataset) for i in range(5): print(tokenized_dataset[i]) training_args = TrainingArguments( output_dir = "***", overwrite_output_dir=True, num_train_epochs=2, gradient_accumulation_steps=1, fp16=True, bf16=False, per_device_train_batch_size=1, optim="paged_adamw_8bit", save_steps=200, logging_steps=25, save_total_limit=2, max_grad_norm=0.3, max_steps=-1, group_by_length=True, learning_rate=2e-3, weight_decay=0.01, warmup_steps=10, lr_scheduler_type="linear", ) peft_params = LoraConfig(lora_alpha=16, lora_dropout=0.1, r=64, bias="none", task_type="CAUSAL_LM") gc.collect() torch.cuda.empty_cache() #4-Bit Quantization( use this if you want to quantize the model using PEFT) os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "max_split_size_mb:40000" trainer = SFTTrainer(model=model, train_dataset=tokenized_dataset, peft_config=peft_params, tokenizer=tokenizer, args=training_args, dataset_text_field="text") trainer.train() trainer.model.save_pretrained(new_model) trainer.tokenizer.save_pretrained(new_model) gc.collect() torch.cuda.empty_cache()
问题根源分析
- 模型未启用4bit量化:代码中定义了
bnb_config但加载模型时未传入,导致Llama-2-7b以全精度(FP32)加载,单卡需约28GB显存,远超单块RTX4090的24GB,直接触发OOM。 - 变量未定义:
updated_dataset和tokenize_function未实现,虽不直接导致OOM,但会引发运行错误,需补充代码。 - 未启用多卡分布式训练:虽指定了
CUDA_VISIBLE_DEVICES="0,1",但未配置分布式参数,模型默认仅加载到单卡,未利用第二块GPU的显存。 - 训练参数不合理:学习率
2e-3过高(Lora微调常规值为2e-4左右),gradient_accumulation_steps=1未充分利用显存空间。
解决方案
- 启用4bit量化并分配多卡:修改模型加载代码,传入量化配置并自动分配到多卡:
model = AutoModelForCausalLM.from_pretrained( base_model, quantization_config=bnb_config, device_map="auto", trust_remote_code=True ) model.config.use_cache = False # 训练时关闭缓存,减少显存占用 model.config.pretraining_tp = 1
- 补充未定义变量:实现分词函数并转换数据集格式:
# 实现分词函数,根据需求调整max_length def tokenize_function(examples): return tokenizer(examples["text"], truncation=True, max_length=512) # 将DataFrame转为Dataset格式 updated_dataset = Dataset.from_pandas(json_dataset)
- 配置分布式训练:在
TrainingArguments中添加分布式参数:
training_args = TrainingArguments( # 原有参数保留 ddp_find_unused_parameters=False, # 启用分布式数据并行 report_to="none", # 关闭第三方报告工具,减少内存占用 )
- 优化训练参数:
- 降低学习率至
2e-4,符合Lora微调常规设置 - 调整
gradient_accumulation_steps=4,等效增大batch size同时降低显存峰值 - 若仍有压力,可进一步减小
max_length的值
- 额外显存优化:
- 保持
optim="paged_adamw_8bit",使用分页优化器减少显存占用 - 训练前执行
gc.collect()和torch.cuda.empty_cache()清理显存 - 训练过程中避免打印大量数据,减少内存消耗
内容的提问来源于stack exchange,提问作者Bansi
相关产品推荐
相关产品推荐

