You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Hugging Face Trainer微调时训练正常但验证阶段CUDA内存不足

大模型微调验证阶段CUDA内存不足问题

用Hugging Face Trainer进行大模型微调时,训练过程可正常执行,但进入验证阶段时触发CUDA内存不足错误。即使将eval_accumulation_steps设置为1也无法解决,参考相关方案后仍未起效。移除TrainingArguments中的验证数据集后程序可正常运行,但添加回验证数据集后,训练至第10步(因eval_steps=10,即将执行验证)时再次出现内存不足错误。

复现代码

import os
os.environ["CUDA_VISIBLE_DEVICES"]="0"
import torch
import torch.nn as nn
import bitsandbytes as bnb
from transformers import AutoTokenizer, AutoConfig, AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained(
    'bigscience/bloom-1b1',
    load_in_8bit=True,
    device_map='auto',
)

tokenizer = AutoTokenizer.from_pretrained('bigscience/bloom-1b1')

from peft import LoraConfig, get_peft_model

config = LoraConfig(
    r= 8, #attention heads
    lora_alpha=32, #alpha scaling
    target_modules=["query_key_value"],
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM" # set this for CLM or Seq2Seq
)

model = get_peft_model(model, config)

import transformers
trainer = transformers.Trainer(
    model=model,
    train_dataset=tokenized_datasets['train'],
    eval_dataset=tokenized_datasets["validation"],
    args=transformers.TrainingArguments(
        per_device_train_batch_size=1,
        gradient_accumulation_steps=4,
        warmup_steps=2,
        max_steps=60,
        learning_rate=2e-4,
        evaluation_strategy = 'steps',
        eval_accumulation_steps = 1, 
        eval_steps = 10,
        seed =  42,
        report_to="wandb",
        fp16=True,
        logging_steps=1,
        output_dir='outputs'
    ),
    data_collator=transformers.DataCollatorForSeq2Seq(tokenizer, pad_to_multiple_of=8, return_tensors="pt", padding=True)

    # data_collator=transformers.DataCollatorForLanguageModeling(tokenizer, mlm=False)
)
model.config.use_cache = False  # silence the warnings. Please re-enable for inference!
trainer.train()

报错信息

OutOfMemoryError                          Traceback (most recent call last)
 in <cell line: 26>()
24 )
25 model.config.use_cache = False  # silence the warnings. Please re-enable for inference!
---> 26 trainer.train()
27
28 wandb.finish()

17 frames
/usr/local/lib/python3.10/dist-packages/torch/nn/functional.py in cross_entropy(input, target, weight, size_average, ignore_index, reduce, reduction, label_smoothing)
3027     if size_average is not None or reduce is not None:
3028         reduction = _Reduction.legacy_get_string(size_average, reduce)
---> 3029     return torch._C._nn.cross_entropy_loss(input, target, weight, _Reduction.get_enum(reduction), ignore_index, label_smoothing)
3030
3031

OutOfMemoryError: CUDA out of memory. Tried to allocate 5.56 GiB (GPU 0; 14.75 GiB total capacity; 12.58 GiB already allocated; 840.81 MiB free; 12.86 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation.  See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF

解决思路

  • 替换数据收集器:当前任务是因果语言模型(CAUSAL_LM),但使用了DataCollatorForSeq2Seq,会生成额外张量(如decoder_input_ids)占用内存。注释掉当前数据收集器,启用代码中已注释的DataCollatorForLanguageModeling。
  • 单独设置验证批次大小:在TrainingArguments中添加per_device_eval_batch_size=1,明确缩小验证阶段的批次规模,进一步降低内存占用。
  • 验证前清理CUDA缓存:自定义Trainer回调,在验证前主动清理缓存:
    from transformers import TrainerCallback
    
    class ClearCacheCallback(TrainerCallback):
        def on_evaluate(self, args, state, control, **kwargs):
            torch.cuda.empty_cache()
    
    trainer.add_callback(ClearCacheCallback())
    
  • 缩小验证数据集规模:先取验证集前几条数据(如前10条)测试,排查是否因验证集过大导致内存溢出。
  • 检查FP16兼容性:确保模型所有组件都兼容FP16,若GPU支持,可尝试替换为bf16=True;也可临时关闭FP16测试,排查是否为精度设置导致的内存问题。
  • 强制验证时禁用use_cache:在训练前或回调中再次设置model.config.use_cache = False,避免验证阶段缓存占用额外内存。

内容的提问来源于stack exchange,提问作者Tommy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 04:42:02