You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch Lightning运行30亿参数T5模型CUDA显存不足问题排查

在4张16GB V100 GPU上运行30亿参数T5模型时遭遇CUDA显存不足问题

我在配备4张16GB显存V100 GPU的节点上运行30亿参数的Transformers T5模型(google/flan-t5-xl),遇到了torch.cuda.OutOfMemoryError: CUDA out of memory错误。使用Lightning框架管理显存,但仍无法定位问题,几乎尝试了Trainer所有DeepSpeed策略参数都没解决,确认硬件应该能支撑该模型运行,希望有人帮忙排查。

疑问:是否需要使用deepspeed.checkpointing.checkpoint?不清楚当前场景下该如何使用。

可复现问题的极简代码

#%% libraries
import torch
import lightning
import transformers
import deepspeed

from torch.optim import AdamW
from torch.utils.data import Dataset, DataLoader
from lightning.pytorch import Trainer, LightningModule, LightningDataModule
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer


#%% data class
class MyDataset(Dataset):
    def __init__(self, checkpoint):
        
        self.tokenizer = AutoTokenizer.from_pretrained(checkpoint)

        # fake data
        self.source_string = 'Apollo 10 (May 18–26, 1969) was a human spaceflight, the fourth crewed mission in the Apollo program of the United States, and the second (after Apollo 8) to orbit the Moon. NASA described it as a "dress rehearsal" for the first Moon landing[4] and designated it an "F" mission, intended to test all spacecraft components and procedures short of actual descent and landing. While astronaut John Young remained in the Command and Service Module (CSM) orbiting the Moon, astronauts Thomas Stafford and Gene Cernan flew the Apollo Lunar Module (LM) to within 14.4 kilometers (7.8 nmi) of the lunar surface, the point at which powered descent for landing would begin on a landing mission, before rejoining Young in the CSM. After orbiting the Moon 31 times, Apollo 10 returned safely to Earth; its success enabled the first crewed landing during the Apollo 11 mission two months later.'
        self.target_string = 'While NASA had considered attempting the first crewed lunar landing on Apollo 10, mission planners ultimately decided that it would be prudent to have a practice flight to hone the procedures and techniques. The crew encountered some problems during the flight: pogo oscillations during the launch phase and a brief, uncontrolled tumble of the LM ascent stage in lunar orbit during its solo flight. However, the mission accomplished its major objectives. Stafford and Cernan observed and photographed Apollo 11s planned landing site in the Sea of Tranquility. Apollo 10 spent 61 hours and 37 minutes orbiting the Moon, for about eight hours of which Stafford and Cernan flew the LM apart from Young in the CSM, and about eight days total in space. Additionally, Apollo 10 set the record for the highest speed attained by a crewed vehicle: 39,897 km/h (11.08 km/s or 24,791 mph) on May 26, 1969, during the return from the Moon.'
        
        self.source_strings = [self.source_string] * 100
        self.target_strings = [self.target_string] * 100
        
        self.source_tokens = [self.tokenizer(el, return_tensors = 'pt') for el in self.source_strings]
        self.target_tokens = [self.tokenizer(el, return_tensors = 'pt') for el in self.target_strings]

    def __len__(self):
        return len(self.source_strings)
    
    def __getitem__(self, idx):
        return self.source_tokens[idx], self.target_tokens[idx]   

#%% lightning module
class MyLightningModule(LightningModule):
    def __init__(self, checkpoint):
        super().__init__()

        self.language_model = AutoModelForSeq2SeqLM.from_pretrained(checkpoint)

    def training_step(self, batch, batch_idx):   
            source = batch[0]
            labels = batch[1].input_ids
            
            outputs = self.language_model(**source, labels = labels)
            #outputs = deepspeed.checkpointing.checkpoint(self.language_model, source.input_ids, source.attention_mask, labels)

            loss, logits = outputs[:2]  
            
            return loss

    def configure_optimizers(self):
        optimizer = deepspeed.ops.adam.FusedAdam(self.parameters(), lr = 1e-5)
        # optimizer = deepspeed.ops.adam.DeepSpeedCPUAdam(self.parameters(), lr = 1e-5)

        return optimizer

#%% instantiate
checkpoint = 'google/flan-t5-xl'
# checkpoint = 'google/flan-t5-base'

train_data = MyDataset(checkpoint)
train_loader = DataLoader(train_data, 
                          batch_size = 1, 
                          collate_fn = lambda x: x[0])

lightning_module = MyLightningModule(checkpoint)

trainer = Trainer(max_epochs = 2,
                  accumulate_grad_batches = 1,
                  accelerator = 'gpu',
                  devices = 4, 
                  strategy = 'deepspeed_stage_3',
                  enable_checkpointing = False,
                  precision = '32',
                 ) 
#%% run
trainer.fit(lightning_module, 
            train_dataloaders = train_loader)

已尝试的解决方案

  • Trainer的strategy参数:deepspeed_stage_2、deepspeed_stage_2_cpu_offload、deepspeed_stage_3、deepspeed_stage_3_cpu_offload、fsdp、fsdp_cpu_offload
  • 使用CPU offload策略时,同步将优化器切换为deepspeed.ops.adam.DeepSpeedCPUAdam

补充说明

当batch size设为1、精度设为16并使用SGD时,模型可以正常运行,但我希望能使用DeepSpeed提供的ADAM优化器。

内容的提问来源于stack exchange,提问作者BioBroo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 03:33:27