You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用HuggingFace Trainer多GPU训练时如何解决内存失衡与OOM错误?

问题

在Azure ML Studio的Standard_NC96ads_A100_v4计算实例(配备4张80GB A100 GPU、96核CPU、880GB系统内存)上,使用HuggingFace Seq2Seq Trainer微调Google的flan-t5-large模型(仅783M参数,单GPU即可容纳),设置单设备训练批次大小为2,训练未执行评估就崩溃。

Seq2Seq训练参数如下:

train_args = Seq2SeqTrainingArguments(
    output_dir=args.output_dir,
    overwrite_output_dir=True,
    evaluation_strategy="epoch",
    save_strategy="epoch",
    logging_steps = 1,
    learning_rate=2e-5,
    per_device_train_batch_size=2,
    per_device_eval_batch_size=50,
    weight_decay=0.01,
    save_total_limit=3,
    num_train_epochs=3,
    predict_with_generate=True,
    fp16=False,
    bf16=True,
    push_to_hub=False,
    report_to="wandb"
)

训练处理600-1000条数据(训练集共约10万条)后触发CUDA OOM错误,GPU内存波动极大,某时刻单GPU内存被完全占满,但4张GPU总内存仅使用39%,波动无固定规律,且发生在模型保存、评估等操作之前。尝试DeepSpeed配置后内存耗尽更早,配置如下:

{
  "bf16": {
    "enabled": "auto"
  },
  "optimizer": {
    "type": "AdamW",
    "params": {
      "lr": "auto",
      "betas": "auto",
      "eps": "auto",
      "weight_decay": "auto"
    }
  },
  "scheduler": {
    "type": "WarmupLR",
    "params": {
      "warmup_min_lr": "auto",
      "warmup_max_lr": "auto",
      "warmup_num_steps": "auto"
    }
  },
  "zero_optimization": {
    "stage": 2,
    "offload_optimizer": {
        "device": "cpu",
        "pin_memory": true
    },
    "allgather_partitions": true,
    "allgather_bucket_size": 2e8,
    "overlap_comm": true,
    "reduce_scatter": true,
    "reduce_bucket_size": 2e8,
    "contiguous_gradients": true
  },
  "gradient_accumulation_steps": "auto",
  "gradient_clipping": "auto",
  "steps_per_print": 2000,
  "train_batch_size": "auto",
  "train_micro_batch_size_per_gpu": "auto",
  "wall_clock_breakdown": false
}

错误信息显示内存碎片化并非问题:

torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 4.81 GiB. GPU has a total capacity of 79.15 GiB of which 2.86 GiB is free. Process 27719 has 76.29 GiB memory in use. Of the allocated memory 74.19 GiB is allocated by PyTorch, and 1.59 GiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)

需采取措施平衡GPU内存使用,避免程序崩溃。

解决方案
  • 强制均匀分配模型参数:手动指定device_map="balanced"或device_map="balanced_low_0",让HuggingFace Accelerate自动将模型参数均匀分配到各GPU,避免单GPU负载过高。
  • 降低日志频率:当前logging_steps=1会每步记录日志,频繁操作易引发内存波动,调整为logging_steps=50或更高,减少内存开销。
  • 调整评估相关设置:关闭predict_with_generate=False(待训练稳定后再开启),同时降低per_device_eval_batch_size至10-20,减少评估阶段的预占内存。
  • 优化DeepSpeed配置:将allgather_bucket_size调小至1e8,关闭overlap_comm(部分环境下重叠通信会引发内存异常);若仍有问题,改用Zero Stage 1,减少参数分片带来的通信内存开销。
  • 启用PyTorch内存优化:设置环境变量PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True,优化内存分配逻辑,减少突发内存占用。
  • 限制GPU内存使用率:初始化模型前添加torch.cuda.set_per_process_memory_fraction(0.9, device=device),限制单进程GPU内存使用率为90%,避免进程占满GPU内存导致后续分配失败。
  • 优化数据加载与预处理:确保数据加载器开启pin_memory=True,设置合理的num_workers(如CPU核心数的1/2);添加样本截断逻辑,限制输入输出序列长度,避免长样本导致单步内存暴增。

内容的提问来源于stack exchange,提问作者Owen Burns

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 10:55:56