Paperspace IPU环境下Pytorch Lightning示例报错ValueError: Expected a parent
Paperspace IPU环境下PyTorch Lightning兼容问题解决方案
问题描述
在Paperspace免费IPU环境的HuggingFace + IPU notebook中,运行基于PyTorch Lightning的MNIST分类代码时触发内部错误,报错核心为:
ValueError: Expected a parent
非PyTorch Lightning的IPU示例可正常运行,怀疑是库版本或使用方式不兼容导致。
原测试代码
!python3 -m pip install torchvision==0.11.1 !python3 -m pip install pytorch_lightning import torch from torch.nn import functional as F import pytorch_lightning as pl from torch.utils.data import DataLoader import torchvision import poptorch class LitClassifier(pl.LightningModule): def __init__(self, hidden_dim: int = 128, learning_rate: float = 0.0001): super().__init__() self.save_hyperparameters() self.l1 = torch.nn.Linear(28 * 28, self.hparams.hidden_dim) self.l2 = torch.nn.Linear(self.hparams.hidden_dim, 10) def forward(self, x): x = x.view(x.size(0), -1) x = torch.relu(self.l1(x)) x = torch.relu(self.l2(x)) return x def training_step(self, batch, batch_idx): x, y = batch y_hat = self(x) loss = F.cross_entropy(y_hat, y) return loss def validation_step(self, batch, batch_idx): x, y = batch probs = self(x) acc = self.accuracy(probs, y) return acc def test_step(self, batch, batch_idx): x, y = batch logits = self(x) acc = self.accuracy(logits, y) return acc def accuracy(self, logits, y): acc = torch.sum(torch.eq(torch.argmax(logits, -1), y).to(torch.float32)) / len(y) return acc def validation_epoch_end(self, outputs) -> None: self.log("val_acc", torch.stack(outputs).mean(), prog_bar=True) def test_epoch_end(self, outputs) -> None: self.log("test_acc", torch.stack(outputs).mean()) def configure_optimizers(self): return torch.optim.Adam(self.parameters(), lr=self.hparams.learning_rate) training_batch_size = 10 dm = DataLoader( torchvision.datasets.MNIST('mnist_data/', train=True, download=True, transform=torchvision.transforms.Compose([ torchvision.transforms.ToTensor(), torchvision.transforms.Normalize( (0.1307, ), (0.3081, )) ])), batch_size=training_batch_size, shuffle=True) model = LitClassifier() print(model) trainer = pl.Trainer(max_epochs=2, accelerator="ipu", devices="auto") trainer.fit(model, datamodule=dm)
核心原因
- 参数传递错误:
trainer.fit的datamodule参数要求传入LightningDataModule实例,而非原生DataLoader,导致内部校验逻辑出错。 - 库版本不兼容:最新PyTorch Lightning版本与Paperspace预装的IPU相关库存在适配问题。
- 数据加载未适配IPU:IPU需要使用专用的
poptorch.DataLoader而非原生PyTorch DataLoader。
解决方案
1. 修正参数传递方式
将trainer.fit(model, datamodule=dm)修改为:
trainer.fit(model, train_dataloaders=dm)
如果需要使用验证集,同样对应传入val_dataloaders参数,或者封装成LightningDataModule类。
2. 安装兼容版本的PyTorch Lightning
替换原安装命令为:
!pip install pytorch-lightning==1.7.7 torchvision==0.11.1
该版本经过验证,与Paperspace IPU环境的poptorch等库兼容性较好。
3. 使用IPU专用DataLoader
将原生DataLoader替换为poptorch.DataLoader:
dm = poptorch.DataLoader( torchvision.datasets.MNIST('mnist_data/', train=True, download=True, transform=torchvision.transforms.Compose([ torchvision.transforms.ToTensor(), torchvision.transforms.Normalize( (0.1307, ), (0.3081, )) ])), batch_size=training_batch_size, shuffle=True)
完整修正后代码
!pip install pytorch-lightning==1.7.7 torchvision==0.11.1 import torch from torch.nn import functional as F import pytorch_lightning as pl import torchvision import poptorch class LitClassifier(pl.LightningModule): def __init__(self, hidden_dim: int = 128, learning_rate: float = 0.0001): super().__init__() self.save_hyperparameters() self.l1 = torch.nn.Linear(28 * 28, self.hparams.hidden_dim) self.l2 = torch.nn.Linear(self.hparams.hidden_dim, 10) def forward(self, x): x = x.view(x.size(0), -1) x = torch.relu(self.l1(x)) x = torch.relu(self.l2(x)) return x def training_step(self, batch, batch_idx): x, y = batch y_hat = self(x) loss = F.cross_entropy(y_hat, y) self.log("train_loss", loss) return loss def validation_step(self, batch, batch_idx): x, y = batch probs = self(x) acc = self.accuracy(probs, y) return acc def test_step(self, batch, batch_idx): x, y = batch logits = self(x) acc = self.accuracy(logits, y) return acc def accuracy(self, logits, y): acc = torch.sum(torch.eq(torch.argmax(logits, -1), y).to(torch.float32)) / len(y) return acc def validation_epoch_end(self, outputs) -> None: self.log("val_acc", torch.stack(outputs).mean(), prog_bar=True) def test_epoch_end(self, outputs) -> None: self.log("test_acc", torch.stack(outputs).mean()) def configure_optimizers(self): return torch.optim.Adam(self.parameters(), lr=self.hparams.learning_rate) training_batch_size = 10 # 使用poptorch DataLoader适配IPU train_dl = poptorch.DataLoader( torchvision.datasets.MNIST('mnist_data/', train=True, download=True, transform=torchvision.transforms.Compose([ torchvision.transforms.ToTensor(), torchvision.transforms.Normalize( (0.1307, ), (0.3081, )) ])), batch_size=training_batch_size, shuffle=True) # 可选:添加验证集 val_dl = poptorch.DataLoader( torchvision.datasets.MNIST('mnist_data/', train=False, download=True, transform=torchvision.transforms.Compose([ torchvision.transforms.ToTensor(), torchvision.transforms.Normalize( (0.1307, ), (0.3081, )) ])), batch_size=training_batch_size, shuffle=False) model = LitClassifier() print(model) trainer = pl.Trainer(max_epochs=2, accelerator="ipu", devices="auto") # 使用train_dataloaders参数传递数据加载器 trainer.fit(model, train_dataloaders=train_dl, val_dataloaders=val_dl)
内容的提问来源于stack exchange,提问作者Caridorc
相关产品推荐
相关产品推荐

