You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

语言模型微调二次迭代时CUDA显存不足问题求助

语言模型微调第二次迭代CUDA内存不足问题

问题详情

在语言模型微调的第二次迭代过程中触发CUDA内存不足错误,代码及报错信息如下:

运行代码

optimizer = AdamW(model.parameters(), lr=1e-4)
scheduler = get_linear_schedule_with_warmup(optimizer, 
                                            num_warmup_steps=steps_per_epoch*1, 
                                            num_training_steps=steps_per_epoch*NUM_EPOCHS)

print("======================= Start pretraining ==============================")

pretrain(model=model,
         train_iter=train_iter,
         valid_iter=valid_iter,
         optimizer=optimizer,
         scheduler=scheduler,
         num_epochs=NUM_EPOCHS)


NUM_EPOCHS = 12
print("======================= Start training ==================================")
optimizer = AdamW(model.parameters(), lr=2e-6)
scheduler = get_linear_schedule_with_warmup(optimizer, 
                                            num_warmup_steps=steps_per_epoch*2, 
                                            num_training_steps=steps_per_epoch*NUM_EPOCHS)

train(model=model, 
      train_iter=train_iter, 
      valid_iter=valid_iter, 
      optimizer=optimizer, 
      scheduler=scheduler, 
      num_epochs=NUM_EPOCHS)

报错信息

torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 16.00 MiB. GPU 0 has a total capacty of 2.00 GiB of which 0 bytes is free.

已尝试调整batch size、梯度裁剪、限制序列长度、调用torch.cuda.empty_cache()、修改优化器参数及epoch数量,但问题仍未解决,需有效解决方案。

可行解决方案

1. 启用梯度累积

若已将batch size调到最小,试试梯度累积——把多个小batch的梯度累加后再更新参数,等效于增大batch size但不额外占用显存。示例代码:

# 设置累积步数,根据显存情况调整
accumulation_steps = 4

for epoch in range(NUM_EPOCHS):
    model.train()
    total_loss = 0.0
    for step, batch in enumerate(train_iter):
        # 前向传播计算损失
        outputs = model(**batch)
        loss = outputs.loss
        # 梯度缩放,避免累积时数值溢出
        loss = loss / accumulation_steps
        loss.backward()
        
        # 累积到指定步数再执行参数更新
        if (step + 1) % accumulation_steps == 0:
            optimizer.step()
            scheduler.step()
            optimizer.zero_grad()
        
        total_loss += loss.item() * accumulation_steps

2. 开启混合精度训练

用PyTorch的自动混合精度(AMP),将部分张量以半精度(FP16)存储,能大幅降低显存占用,同时几乎不影响模型性能:

from torch.cuda.amp import GradScaler, autocast

# 初始化混合精度缩放器
scaler = GradScaler()

for epoch in range(NUM_EPOCHS):
    model.train()
    for batch in train_iter:
        optimizer.zero_grad()
        # 自动混合精度上下文,自动处理精度转换
        with autocast():
            outputs = model(**batch)
            loss = outputs.loss
        # 缩放梯度后反向传播
        scaler.scale(loss).backward()
        scaler.step(optimizer)
        scaler.update()
        scheduler.step()

3. 清理预训练阶段的冗余显存

代码中先执行pretrain再启动train,预训练后可能残留未释放的中间张量、优化器或调度器缓存。可以在预训练结束后手动清理:

# 完成预训练后
pretrain(...)
# 删除预训练用的优化器和调度器
del optimizer, scheduler
# 强制清理CUDA缓存
torch.cuda.empty_cache()
# 重新初始化训练阶段的优化器和调度器
optimizer = AdamW(model.parameters(), lr=2e-6)
scheduler = get_linear_schedule_with_warmup(optimizer, 
                                            num_warmup_steps=steps_per_epoch*2, 
                                            num_training_steps=steps_per_epoch*NUM_EPOCHS)

4. 压缩模型显存占用

如果模型本身超出2GB显存承载能力,试试这些方法:

  • 替换为轻量级模型变体,比如用DistilBERT替代BERT-base,TinyBERT替代DistilBERT
  • 启用梯度检查点,牺牲少量计算速度换取显存空间(以Hugging Face模型为例):
model.gradient_checkpointing_enable()

5. 排查数据加载器的显存泄漏

确保数据加载器没有提前将数据加载到GPU,或存在重复张量引用。可以在每个batch处理完后手动清理:

for batch in train_iter:
    # 处理当前batch的训练逻辑
    ...
    # 删除batch张量并清理缓存
    del batch
    torch.cuda.empty_cache()

内容的提问来源于stack exchange,提问作者Geanry Blog

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 21:43:20