PyTorch Lightning运行30亿参数T5模型CUDA显存不足问题排查
在4张16GB V100 GPU上运行30亿参数T5模型时遭遇CUDA显存不足问题
我在配备4张16GB显存V100 GPU的节点上运行30亿参数的Transformers T5模型(google/flan-t5-xl),遇到了torch.cuda.OutOfMemoryError: CUDA out of memory错误。使用Lightning框架管理显存,但仍无法定位问题,几乎尝试了Trainer所有DeepSpeed策略参数都没解决,确认硬件应该能支撑该模型运行,希望有人帮忙排查。
疑问:是否需要使用deepspeed.checkpointing.checkpoint?不清楚当前场景下该如何使用。
可复现问题的极简代码
#%% libraries import torch import lightning import transformers import deepspeed from torch.optim import AdamW from torch.utils.data import Dataset, DataLoader from lightning.pytorch import Trainer, LightningModule, LightningDataModule from transformers import AutoModelForSeq2SeqLM, AutoTokenizer #%% data class class MyDataset(Dataset): def __init__(self, checkpoint): self.tokenizer = AutoTokenizer.from_pretrained(checkpoint) # fake data self.source_string = 'Apollo 10 (May 18–26, 1969) was a human spaceflight, the fourth crewed mission in the Apollo program of the United States, and the second (after Apollo 8) to orbit the Moon. NASA described it as a "dress rehearsal" for the first Moon landing[4] and designated it an "F" mission, intended to test all spacecraft components and procedures short of actual descent and landing. While astronaut John Young remained in the Command and Service Module (CSM) orbiting the Moon, astronauts Thomas Stafford and Gene Cernan flew the Apollo Lunar Module (LM) to within 14.4 kilometers (7.8 nmi) of the lunar surface, the point at which powered descent for landing would begin on a landing mission, before rejoining Young in the CSM. After orbiting the Moon 31 times, Apollo 10 returned safely to Earth; its success enabled the first crewed landing during the Apollo 11 mission two months later.' self.target_string = 'While NASA had considered attempting the first crewed lunar landing on Apollo 10, mission planners ultimately decided that it would be prudent to have a practice flight to hone the procedures and techniques. The crew encountered some problems during the flight: pogo oscillations during the launch phase and a brief, uncontrolled tumble of the LM ascent stage in lunar orbit during its solo flight. However, the mission accomplished its major objectives. Stafford and Cernan observed and photographed Apollo 11s planned landing site in the Sea of Tranquility. Apollo 10 spent 61 hours and 37 minutes orbiting the Moon, for about eight hours of which Stafford and Cernan flew the LM apart from Young in the CSM, and about eight days total in space. Additionally, Apollo 10 set the record for the highest speed attained by a crewed vehicle: 39,897 km/h (11.08 km/s or 24,791 mph) on May 26, 1969, during the return from the Moon.' self.source_strings = [self.source_string] * 100 self.target_strings = [self.target_string] * 100 self.source_tokens = [self.tokenizer(el, return_tensors = 'pt') for el in self.source_strings] self.target_tokens = [self.tokenizer(el, return_tensors = 'pt') for el in self.target_strings] def __len__(self): return len(self.source_strings) def __getitem__(self, idx): return self.source_tokens[idx], self.target_tokens[idx] #%% lightning module class MyLightningModule(LightningModule): def __init__(self, checkpoint): super().__init__() self.language_model = AutoModelForSeq2SeqLM.from_pretrained(checkpoint) def training_step(self, batch, batch_idx): source = batch[0] labels = batch[1].input_ids outputs = self.language_model(**source, labels = labels) #outputs = deepspeed.checkpointing.checkpoint(self.language_model, source.input_ids, source.attention_mask, labels) loss, logits = outputs[:2] return loss def configure_optimizers(self): optimizer = deepspeed.ops.adam.FusedAdam(self.parameters(), lr = 1e-5) # optimizer = deepspeed.ops.adam.DeepSpeedCPUAdam(self.parameters(), lr = 1e-5) return optimizer #%% instantiate checkpoint = 'google/flan-t5-xl' # checkpoint = 'google/flan-t5-base' train_data = MyDataset(checkpoint) train_loader = DataLoader(train_data, batch_size = 1, collate_fn = lambda x: x[0]) lightning_module = MyLightningModule(checkpoint) trainer = Trainer(max_epochs = 2, accumulate_grad_batches = 1, accelerator = 'gpu', devices = 4, strategy = 'deepspeed_stage_3', enable_checkpointing = False, precision = '32', ) #%% run trainer.fit(lightning_module, train_dataloaders = train_loader)
已尝试的解决方案
- Trainer的strategy参数:
deepspeed_stage_2、deepspeed_stage_2_cpu_offload、deepspeed_stage_3、deepspeed_stage_3_cpu_offload、fsdp、fsdp_cpu_offload - 使用CPU offload策略时,同步将优化器切换为
deepspeed.ops.adam.DeepSpeedCPUAdam
补充说明
当batch size设为1、精度设为16并使用SGD时,模型可以正常运行,但我希望能使用DeepSpeed提供的ADAM优化器。
内容的提问来源于stack exchange,提问作者BioBroo
相关产品推荐
相关产品推荐

