You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

训练正常但验证阶段触发CUDA Out Of Memory Error求助

T5-small拼写纠错模型验证阶段CUDA内存不足问题解决

问题背景

在Colab T4 GPU上基于T5-small训练拼写纠错模型,初始用10k条评论训练触发CUDA内存不足,缩减到2k条后训练成功并保存模型,但验证阶段仍报相同内存错误,试过重启会话、调大gradient_accumulation_steps均无效。

报错信息

CUDA out of memory. Tried to allocate 6.99 GiB. GPU 0 has a total capacity of 14.74 GiB of which 6.95 GiB is free. Process 79643 has 7.79 GiB memory in use. Of the allocated memory 7.59 GiB is allocated by PyTorch, and 76.39 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management...

用户代码

os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
from transformers import T5ForConditionalGeneration, T5Tokenizer, Trainer, TrainingArguments, DataCollatorForSeq2Seq
from datasets import load_dataset

# Load dataset
dataset = load_dataset("csv", data_files="elaborate_cosmetic_reviews.csv")

# load the pre-trained Tokenizer and Model
fine_model_name = "./spellcheck_model"
tokenizer = T5Tokenizer.from_pretrained(fine_model_name, local_files_only=True)
model = T5ForConditionalGeneration.from_pretrained(fine_model_name, local_files_only=True)

# Define training arguments
training_args = TrainingArguments(
    output_dir="./spellcheck_model_base",
    per_device_train_batch_size=2,
    report_to="none",
    fp16=True, 
    gradient_accumulation_steps = 256,
)

# Create data collator
data_collator = DataCollatorForSeq2Seq(tokenizer, model=model)

trainer = Trainer(
    model=model,
    args=training_args,
    data_collator=data_collator,
    compute_metrics=compute_metrics,
    eval_dataset=tokenized_train, # Use tokenized_eval for test set
)
import torch
torch.cuda.empty_cache()

# Evaluate on the training set
print("Training set Evaluation: ")
train_metrics = trainer.evaluate(eval_dataset=tokenized_train)
print(train_metrics)

解决方法

  • 单独设置验证批量大小:T5生成阶段(验证时要输出纠错文本)比训练阶段更耗内存,TrainingArguments默认复用训练批量大小,直接设置per_device_eval_batch_size=1,进一步降低单步内存占用。
  • 提前设置CUDA环境变量:当前环境变量设置在导入transformers之后,PyTorch已经初始化CUDA,设置不会生效。必须把环境变量放在所有库导入之前。
  • 启用验证阶段混合精度:训练时开了fp16=True,但验证阶段默认可能没启用,加上fp16_full_eval=True,强制用FP16推理,减少内存消耗。
  • 提前清理GPU缓存:torch.cuda.empty_cache()要放在模型加载之前,加载模型前释放残留内存,而不是加载后调用。
  • 缩小验证集规模:如果只是快速验证效果,不用全量2k数据,随机抽几百条即可,比如tokenized_train.select(range(500))作为验证集。

修改后的示例代码

import os
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"

from transformers import T5ForConditionalGeneration, T5Tokenizer, Trainer, TrainingArguments, DataCollatorForSeq2Seq
from datasets import load_dataset
import torch

# 提前清理GPU缓存
torch.cuda.empty_cache()

# Load dataset
dataset = load_dataset("csv", data_files="elaborate_cosmetic_reviews.csv")

# load the pre-trained Tokenizer and Model
fine_model_name = "./spellcheck_model"
tokenizer = T5Tokenizer.from_pretrained(fine_model_name, local_files_only=True)
model = T5ForConditionalGeneration.from_pretrained(fine_model_name, local_files_only=True)

# Define training arguments
training_args = TrainingArguments(
    output_dir="./spellcheck_model_base",
    per_device_train_batch_size=2,
    per_device_eval_batch_size=1,  # 单独设置验证批量
    report_to="none",
    fp16=True, 
    fp16_full_eval=True,  # 验证阶段启用混合精度
    gradient_accumulation_steps = 256,
)

# Create data collator
data_collator = DataCollatorForSeq2Seq(tokenizer, model=model)

# 缩小验证集规模,只取前500条
small_eval_dataset = tokenized_train.select(range(500))

trainer = Trainer(
    model=model,
    args=training_args,
    data_collator=data_collator,
    compute_metrics=compute_metrics,
    eval_dataset=small_eval_dataset,
)

# Evaluate on the training set
print("Training set Evaluation: ")
train_metrics = trainer.evaluate()
print(train_metrics)

内容的提问来源于stack exchange,提问作者Anurag Pandey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 07:57:27