Azure ML Studio中PyTorch报错:'NoneType'对象无state_dict属性
解决Azure ML中PyTorch Lightning保存Checkpoint时的AttributeError问题
问题现象
代码在Google Colab运行正常,但在Azure ML Notebook中执行时出现以下错误:
File /anaconda/envs/azureml_py38_PT_TF/lib/python3.8/site-packages/pytorch_lightning/trainer/training_io.py:268, in TrainerIOMixin.save_checkpoint(self, filepath, weights_only) 267 def save_checkpoint(self, filepath, weights_only: bool = False): --> 268 checkpoint = self.dump_checkpoint(weights_only) 270 if self.is_global_zero: 271 # do the actual save 272 try: File /anaconda/envs/azureml_py38_PT_TF/lib/python3.8/site-packages/pytorch_lightning/trainer/training_io.py:362, in TrainerIOMixin.dump_checkpoint(self, weights_only) 360 # save native amp scaling 361 if self.use_amp and NATIVE_AMP_AVALAIBLE and not self.use_tpu: --> 362 checkpoint['native_amp_scaling_state'] = self.scaler.state_dict() 364 # add the module_arguments and state_dict from the model 365 model = self.get_model() AttributeError: 'NoneType' object has no attribute 'state_dict'
使用的模型代码如下:
import torch import torch.nn as nn import torch.nn.functional as F from torch.utils.data import DataLoader from collections import OrderedDict import pytorch_lightning as pl import time # 假设EvaluationDataset和LABEL_COUNT已定义 class EvaluationModel(pl.LightningModule): def __init__(self,learning_rate=1e-3,batch_size=1024,layer_count=10): super().__init__() self.batch_size = batch_size self.learning_rate = learning_rate layers = [] for i in range(layer_count-1): layers.append((f"linear-{i}", nn.Linear(808, 808))) layers.append((f"relu-{i}", nn.ReLU())) layers.append((f"linear-{layer_count-1}", nn.Linear(808, 1))) self.seq = nn.Sequential(OrderedDict(layers)) def forward(self, x): return self.seq(x) def training_step(self, batch, batch_idx): x, y = batch['binary'], batch['eval'] y_hat = self(x) loss = F.l1_loss(y_hat, y) self.log("train_loss", loss) return loss def configure_optimizers(self): return torch.optim.Adam(self.parameters(), lr=self.learning_rate) def train_dataloader(self): dataset = EvaluationDataset(count=LABEL_COUNT) return DataLoader(dataset, batch_size=self.batch_size, num_workers=2, pin_memory=True) configs = [ {"layer_count": 4, "batch_size": 512}, # {"layer_count": 6, "batch_size": 1024}, ] for config in configs: version_name = f'{int(time.time())}-batch_size-{config["batch_size"]}-layer_count-{config["layer_count"]}' logger = pl.loggers.TensorBoardLogger("lightning_logs", name="chessml", version=version_name) trainer = pl.Trainer(gpus=1,precision=16,max_epochs=1,auto_lr_find=True,logger=logger) model = EvaluationModel(layer_count=config["layer_count"],batch_size=config["batch_size"],learning_rate=1e-3) # trainer.tune(model) # lr_finder = trainer.tuner.lr_find(model, min_lr=1e-6, max_lr=1e-3, num_training=25) # fig = lr_finder.plot(suggest=True) # fig.show() trainer.fit(model) break
解决方案
错误根源是开启混合精度(precision=16)后,PyTorch Lightning的AMP scaler未被正确初始化,导致self.scaler为None。可通过以下两种方式解决:
方法1:显式指定AMP后端
在初始化Trainer时添加amp_backend='native'参数,强制使用原生AMP并正确初始化scaler:
trainer = pl.Trainer(gpus=1, precision=16, amp_backend='native', max_epochs=1, auto_lr_find=True, logger=logger)
方法2:关闭混合精度
如果不需要混合精度训练,直接将precision=16改为precision=32:
trainer = pl.Trainer(gpus=1, precision=32, max_epochs=1, auto_lr_find=True, logger=logger)
额外建议
检查Azure ML环境中PyTorch Lightning的版本是否与Colab一致,版本差异可能导致AMP初始化逻辑不同。可通过以下命令查看版本:
pip show pytorch-lightning
内容的提问来源于stack exchange,提问作者Marcus O'Yang
相关产品推荐
相关产品推荐

