You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Llama 3微调时CUDA启动失败报错的排查求助

CUDA报错排查与解决(Llama 3 QLoRA微调)

环境与问题描述

使用2块24GB GPU,基于1L数据集微调Llama 3-8B-Instruct模型,训练至30k步时出现如下CUDA报错:

Traceback (most recent call last):
  File "/home/llm04/Ejyle_Sutherland_NLP/finetune/finetune_with_peft.py", line 152, in <module>
    trainer.train()
  File "/home/llm04/new_venv/lib/python3.10/site-packages/trl/trainer/sft_trainer.py", line 440, in train
    output = super().train(*args, **kwargs)
  File "/home/llm04/new_venv/lib/python3.10/site-packages/transformers/trainer.py", line 1885, in train
    return inner_training_loop(
  File "/home/llm04/new_venv/lib/python3.10/site-packages/transformers/trainer.py", line 2216, in _inner_training_loop
    tr_loss_step = self.training_step(model, inputs)
  File "/home/llm04/new_venv/lib/python3.10/site-packages/transformers/trainer.py", line 3241, in training_step
    torch.cuda.empty_cache()
  File "/home/llm04/new_venv/lib/python3.10/site-packages/torch/cuda/memory.py", line 162, in empty_cache
    torch._C._cuda_emptyCache()
RuntimeError: CUDA error: unspecified launch failure
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1.
Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.

已尝试重启服务器、新建虚拟环境,问题仍未解决,使用的微调代码如下:

from transformers import (
    AutoModelForCausalLM,
    AutoTokenizer,
    BitsAndBytesConfig,
    TrainingArguments,
    pipeline,
    logging,
)
from peft import (
    LoraConfig,
    get_peft_model,
)
import os
import torch
from torch.utils.data import DataLoader
import wandb
import pandas as pd
from datasets import Dataset, load_dataset
from trl import SFTTrainer, setup_chat_format
from huggingface_hub import login

torch.cuda.empty_cache()

# # Insert your Hugging Face token here
hf_token = "token"

# # Login to Hugging Face Hub
login(token=hf_token)

# # Define paths and model parameters
# base_model = "meta-llama/Meta-Llama-3-8B"
base_model = "/home/llm04/Meta-Llama-3-8B-Instruct"
dataset_path = "/home/llm04/Ejyle_Sutherland_NLP/Datasets/R918.xlsx"
new_model = "R918-Test-10"
torch_dtype = torch.bfloat16
attn_implementation = "eager"

# QLoRA config
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch_dtype,
    bnb_4bit_use_double_quant=True,
)

# Load model
model = AutoModelForCausalLM.from_pretrained(
    base_model,
    quantization_config=bnb_config,
    device_map="auto",
    attn_implementation=attn_implementation
)

# Load tokenizer
tokenizer = AutoTokenizer.from_pretrained(base_model)
model, tokenizer = setup_chat_format(model, tokenizer)

    

# LoRA config
peft_config = LoraConfig(
    r=16,
    lora_alpha=32,
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM",
    target_modules=['up_proj', 'down_proj', 'gate_proj', 'k_proj', 'q_proj', 'v_proj', 'o_proj'],
)
model = get_peft_model(model, peft_config)
# Load and preprocess the dataset
df = pd.read_excel(dataset_path)
dataset = Dataset.from_pandas(df)
#dataset=load_dataset('csv',data_files=dataset_path)

system_prompt = (
    """you are a medical coding expert, who have access to medical coding guidelines. extract all the medical conditions from all the sections including clinical information/condition and its associated ICD-10 descriptions(description only without code) with anatomical locations for the given clinical text of all the sections, exclude negated conditions from the given radiology record and return only the ICD-10 descriptions(description only without code) of non-negated conditions in this json format{"description":[]}, except json do not send anything else."""
)

template_tokenizer= AutoTokenizer.from_pretrained("HuggingFaceH4/zephyr-7b-beta")

def format_chat_template(row):
    user_content = row["Sentence"] if row["Sentence"] is not None else ""
    assistant_content = row["Description"] if row["Description"] is not None else ""
    messages = [
                {"role": "system", "content": system_prompt},
                {"role": "user", "content": user_content},
                {"role": "assistant", "content": assistant_content}]
    #print(row_json,"------------------------------------------")
    final = template_tokenizer.apply_chat_template(messages, tokenize=False)
    #print(final,"----------")   
    return {'text': final}

#print(format_chat_template(dataset[0]))
print(f"original dataset: {len(dataset)}")
dataset = dataset.map(format_chat_template, remove_columns=dataset.column_names)
print(dataset[0],"---------------------------------map-------------------------------")
print(len(dataset))
dataset = dataset.shuffle(seed=66).select(range(563))

print(f"shuffled dataset: {len(dataset)}")

# Split the dataset into a training and validation set
dataset = dataset.train_test_split(test_size=0.00539)
print(dataset['train'][0],"----------------split-------------------")
# Training parameters
training_arguments = TrainingArguments(
    output_dir=new_model,
    per_device_train_batch_size=1,
    per_device_eval_batch_size=1,
    gradient_accumulation_steps=2,
    optim="paged_adamw_32bit",
    num_train_epochs=5,
    eval_strategy="epoch",
    eval_steps=100,  # Change to 10 to avoid too frequent evaluation
    logging_steps=50,
    warmup_steps=100,
    logging_strategy="epoch",
    learning_rate=2e-5,
    fp16=False,
    bf16=True,
    group_by_length=True,
    load_best_model_at_end=True,
    metric_for_best_model="loss",
    save_total_limit=3,
    save_strategy="epoch",
    #max_steps=(len(dataset)//per_device_train_batch_size)*num_train_epochs
    # report_to="wandb"
)
device = torch.device("cuda:0")
peft_model= model.to(device)
# Supervised fine-tuning (SFT) trainer
trainer = SFTTrainer(
    model=model,
    train_dataset=dataset["train"],
    eval_dataset=dataset["test"],
    peft_config=peft_config,
    dataset_text_field="text",
    tokenizer=tokenizer,
    args=training_arguments,
    max_seq_length=1024,
    packing=False,
    )

trainer.train()


model.config.use_cache = True
trainer.model.save_pretrained(new_model)
# trainer.model.push_to_hub(new_model, use_temp_dir=False)

排查与解决方案

1. 定位真实错误来源

按报错提示,先开启CUDA_LAUNCH_BLOCKING环境变量,强制CUDA同步执行,获取准确的错误栈:

export CUDA_LAUNCH_BLOCKING=1

再启动训练脚本,此时报错会指向真正触发内核失败的代码位置,而非empty_cache环节。

2. 修复设备映射冲突

代码中使用device_map="auto"自动分配多GPU模型参数,之后又手动执行model.to(device),这会破坏自动设备映射,导致模型参数跨设备分布异常,触发CUDA错误。删除以下代码行:

device = torch.device("cuda:0")
peft_model= model.to(device)

3. 优化显存管理,减少碎片化

  • 调高gradient_accumulation_steps:从2改为4或8,降低单步训练的显存占用;
  • 开启梯度检查点:在加载模型时添加gradient_checkpointing=True,或在TrainingArguments中设置gradient_checkpointing=True;
  • 降低max_seq_length:暂时从1024改为512,排查是否因长序列导致显存溢出;
  • 移除训练前的torch.cuda.empty_cache()调用:频繁清空缓存会加剧显存碎片化,transformers框架会自动管理显存。

4. 检查数据集与训练配置

  • 排查格式化后的数据集:检查format_chat_template生成的text字段是否存在空内容、超长文本或格式错误,避免tokenization时出现异常;
  • 关闭group_by_length=True:该配置可能导致部分batch的序列长度远超平均,触发显存不足。

5. 验证依赖版本兼容性

确保以下库的版本匹配:

  • transformers >= 4.38.0
  • trl >= 0.7.0
  • peft >= 0.8.0
  • bitsandbytes >= 0.41.0
    同时确认PyTorch版本与系统CUDA版本兼容(如PyTorch 2.1对应CUDA 11.8/12.1)。

内容的提问来源于stack exchange,提问作者Siddharth S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 06:47:12