Llama 3微调时CUDA启动失败报错的排查求助
CUDA报错排查与解决(Llama 3 QLoRA微调)
环境与问题描述
使用2块24GB GPU,基于1L数据集微调Llama 3-8B-Instruct模型,训练至30k步时出现如下CUDA报错:
Traceback (most recent call last): File "/home/llm04/Ejyle_Sutherland_NLP/finetune/finetune_with_peft.py", line 152, in <module> trainer.train() File "/home/llm04/new_venv/lib/python3.10/site-packages/trl/trainer/sft_trainer.py", line 440, in train output = super().train(*args, **kwargs) File "/home/llm04/new_venv/lib/python3.10/site-packages/transformers/trainer.py", line 1885, in train return inner_training_loop( File "/home/llm04/new_venv/lib/python3.10/site-packages/transformers/trainer.py", line 2216, in _inner_training_loop tr_loss_step = self.training_step(model, inputs) File "/home/llm04/new_venv/lib/python3.10/site-packages/transformers/trainer.py", line 3241, in training_step torch.cuda.empty_cache() File "/home/llm04/new_venv/lib/python3.10/site-packages/torch/cuda/memory.py", line 162, in empty_cache torch._C._cuda_emptyCache() RuntimeError: CUDA error: unspecified launch failure CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1. Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.
已尝试重启服务器、新建虚拟环境,问题仍未解决,使用的微调代码如下:
from transformers import ( AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig, TrainingArguments, pipeline, logging, ) from peft import ( LoraConfig, get_peft_model, ) import os import torch from torch.utils.data import DataLoader import wandb import pandas as pd from datasets import Dataset, load_dataset from trl import SFTTrainer, setup_chat_format from huggingface_hub import login torch.cuda.empty_cache() # # Insert your Hugging Face token here hf_token = "token" # # Login to Hugging Face Hub login(token=hf_token) # # Define paths and model parameters # base_model = "meta-llama/Meta-Llama-3-8B" base_model = "/home/llm04/Meta-Llama-3-8B-Instruct" dataset_path = "/home/llm04/Ejyle_Sutherland_NLP/Datasets/R918.xlsx" new_model = "R918-Test-10" torch_dtype = torch.bfloat16 attn_implementation = "eager" # QLoRA config bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch_dtype, bnb_4bit_use_double_quant=True, ) # Load model model = AutoModelForCausalLM.from_pretrained( base_model, quantization_config=bnb_config, device_map="auto", attn_implementation=attn_implementation ) # Load tokenizer tokenizer = AutoTokenizer.from_pretrained(base_model) model, tokenizer = setup_chat_format(model, tokenizer) # LoRA config peft_config = LoraConfig( r=16, lora_alpha=32, lora_dropout=0.05, bias="none", task_type="CAUSAL_LM", target_modules=['up_proj', 'down_proj', 'gate_proj', 'k_proj', 'q_proj', 'v_proj', 'o_proj'], ) model = get_peft_model(model, peft_config) # Load and preprocess the dataset df = pd.read_excel(dataset_path) dataset = Dataset.from_pandas(df) #dataset=load_dataset('csv',data_files=dataset_path) system_prompt = ( """you are a medical coding expert, who have access to medical coding guidelines. extract all the medical conditions from all the sections including clinical information/condition and its associated ICD-10 descriptions(description only without code) with anatomical locations for the given clinical text of all the sections, exclude negated conditions from the given radiology record and return only the ICD-10 descriptions(description only without code) of non-negated conditions in this json format{"description":[]}, except json do not send anything else.""" ) template_tokenizer= AutoTokenizer.from_pretrained("HuggingFaceH4/zephyr-7b-beta") def format_chat_template(row): user_content = row["Sentence"] if row["Sentence"] is not None else "" assistant_content = row["Description"] if row["Description"] is not None else "" messages = [ {"role": "system", "content": system_prompt}, {"role": "user", "content": user_content}, {"role": "assistant", "content": assistant_content}] #print(row_json,"------------------------------------------") final = template_tokenizer.apply_chat_template(messages, tokenize=False) #print(final,"----------") return {'text': final} #print(format_chat_template(dataset[0])) print(f"original dataset: {len(dataset)}") dataset = dataset.map(format_chat_template, remove_columns=dataset.column_names) print(dataset[0],"---------------------------------map-------------------------------") print(len(dataset)) dataset = dataset.shuffle(seed=66).select(range(563)) print(f"shuffled dataset: {len(dataset)}") # Split the dataset into a training and validation set dataset = dataset.train_test_split(test_size=0.00539) print(dataset['train'][0],"----------------split-------------------") # Training parameters training_arguments = TrainingArguments( output_dir=new_model, per_device_train_batch_size=1, per_device_eval_batch_size=1, gradient_accumulation_steps=2, optim="paged_adamw_32bit", num_train_epochs=5, eval_strategy="epoch", eval_steps=100, # Change to 10 to avoid too frequent evaluation logging_steps=50, warmup_steps=100, logging_strategy="epoch", learning_rate=2e-5, fp16=False, bf16=True, group_by_length=True, load_best_model_at_end=True, metric_for_best_model="loss", save_total_limit=3, save_strategy="epoch", #max_steps=(len(dataset)//per_device_train_batch_size)*num_train_epochs # report_to="wandb" ) device = torch.device("cuda:0") peft_model= model.to(device) # Supervised fine-tuning (SFT) trainer trainer = SFTTrainer( model=model, train_dataset=dataset["train"], eval_dataset=dataset["test"], peft_config=peft_config, dataset_text_field="text", tokenizer=tokenizer, args=training_arguments, max_seq_length=1024, packing=False, ) trainer.train() model.config.use_cache = True trainer.model.save_pretrained(new_model) # trainer.model.push_to_hub(new_model, use_temp_dir=False)
排查与解决方案
1. 定位真实错误来源
按报错提示,先开启CUDA_LAUNCH_BLOCKING环境变量,强制CUDA同步执行,获取准确的错误栈:
export CUDA_LAUNCH_BLOCKING=1
再启动训练脚本,此时报错会指向真正触发内核失败的代码位置,而非empty_cache环节。
2. 修复设备映射冲突
代码中使用device_map="auto"自动分配多GPU模型参数,之后又手动执行model.to(device),这会破坏自动设备映射,导致模型参数跨设备分布异常,触发CUDA错误。删除以下代码行:
device = torch.device("cuda:0") peft_model= model.to(device)
3. 优化显存管理,减少碎片化
- 调高
gradient_accumulation_steps:从2改为4或8,降低单步训练的显存占用; - 开启梯度检查点:在加载模型时添加
gradient_checkpointing=True,或在TrainingArguments中设置gradient_checkpointing=True; - 降低
max_seq_length:暂时从1024改为512,排查是否因长序列导致显存溢出; - 移除训练前的
torch.cuda.empty_cache()调用:频繁清空缓存会加剧显存碎片化,transformers框架会自动管理显存。
4. 检查数据集与训练配置
- 排查格式化后的数据集:检查
format_chat_template生成的text字段是否存在空内容、超长文本或格式错误,避免tokenization时出现异常; - 关闭
group_by_length=True:该配置可能导致部分batch的序列长度远超平均,触发显存不足。
5. 验证依赖版本兼容性
确保以下库的版本匹配:
transformers >= 4.38.0trl >= 0.7.0peft >= 0.8.0bitsandbytes >= 0.41.0
同时确认PyTorch版本与系统CUDA版本兼容(如PyTorch 2.1对应CUDA 11.8/12.1)。
内容的提问来源于stack exchange,提问作者Siddharth S
相关产品推荐
相关产品推荐

