本地GPU训练PyTorch LeNet时Jupyter内核崩溃求助
本地Jupyter用GPU训练LeNet时内核崩溃问题
问题描述
- 实验内容:使用PyTorch实现LeNet神经网络,在本地Jupyter Notebook中尝试通过GPU训练MNIST数据集
- 异常现象:该网络在Google Colab的CPU/GPU运行时、本地CPU环境下均可正常训练,但使用本地GPU时会直接导致Jupyter内核崩溃,仅提示「内核似乎已死亡」,无额外错误日志
- 已排除情况:基础线性网络在本地GPU上可正常运行;系统可用内存16GB(总32GB),GPU为RTX4090,显存充足,排除内存/显存不足问题
LeNet实现代码
# Importing all dependencies import os # for some OS ops import torch from torch import nn torch.manual_seed(0) torch.backends.cudnn.deterministic = True torch.backends.cudnn.benchmark = False from torch.utils.data import DataLoader from torchvision import transforms import torchvision from torchvision.datasets import MNIST # The data set that we will use import matplotlib.pyplot as plt import numpy as np import time class LeNet(nn.Module): def __init__(self, inputSize=28 * 28, outputSize=10, lr=0.01): super().__init__() self.layer1 = nn.Sequential( nn.Conv2d(in_channels=1, out_channels=6, kernel_size=5, padding=2), nn.BatchNorm2d(6), nn.Sigmoid(), nn.AvgPool2d(kernel_size=2, stride=2) ) self.layer2 = nn.Sequential( nn.Conv2d(6, 16, kernel_size=5, padding=0), nn.BatchNorm2d(16), nn.Sigmoid(), nn.AvgPool2d(kernel_size=2, stride=2) ) self.fc = nn.Linear(400, 120) self.relu = nn.ReLU() self.fc1 = nn.Linear(120, 84) self.relu1 = nn.ReLU() self.fc2 = nn.Linear(84, outputSize) # Setting the learning rate self.lr = lr ## The forward step def forward(self, x): # the crash occurs during the next line of code, with no output out = self.layer1(x) out = self.layer2(out) out = out.reshape(out.size(0), -1) out = self.fc(out) out = self.relu(out) out = self.fc1(out) out = self.relu1(out) out = self.fc2(out) return out ## The loss function - Here, we will use Cross Entropy Loss def loss(self, y_hat, y): fn = nn.CrossEntropyLoss() return fn(y_hat, y) ## The optimization algorithm def configure_optimizers(self): return torch.optim.Adam(self.parameters(), self.lr) class Trainer: def __init__(self, n_epochs = 3): self.max_epochs = n_epochs return # The fitting step def fit(self, model, data): self.data = data # configure the optimizer self.optimizer = model.configure_optimizers() self.model = model for epoch in range(self.max_epochs): epoch_start_time = time.time() self.fit_epoch() print("Epoch " + str(epoch) + " finished in " + str(time.time() - epoch_start_time) + "s") print("Training process has finished") def fit_epoch(self): current_loss = 0.0 # iterate over the DataLoader for training data # This iteration is over the batches # For each batch, it updates the network weights and computes the loss for i, data in enumerate(self.data): # Get input and its corresponding groundtruth output inputs, target = data # Clear gradient buffers because we don't want any gradient from previous # epoch to carry forward, dont want to cummulate gradients self.optimizer.zero_grad() # get output from the model, given the inputs # crashes while doing this outputs = self.model(inputs) # get loss for the predicted output loss = self.model.loss(outputs, target) # get gradients w.r.t the parameters of the model loss.backward() # update the parameters (perform optimization) self.optimizer.step() # Let's print some statisics current_loss += loss.item() if i % 500 == 499: print('Loss after mini-batch %5d: %.3f' % (i + 1, current_loss / 500)) current_loss = 0.0 # Training on the GPU def to_device(data, device): # Move tensor(s) to chosen device if isinstance(data, (list,tuple)): return [to_device(x, device) for x in data] return data.to(device, non_blocking=True) class DeviceDataLoader(): # Wrap a dataloader to move data to a device def __init__(self, dl, device): self.dl = dl self.device = device def __iter__(self): # Yield a batch of data after moving it to device for b in self.dl: yield to_device(b, self.device) def __len__(self): # Number of batches return len(self.dl) device = torch.device("cuda:0" if torch.cuda.is_available() else "cpu") #device = torch.device("cpu") # to force it to use CPU print(device) # 1. Loading the MNIST data set # Transforms to apply to the data transform = transforms.Compose([ transforms.Resize((28,28)), # I dont think I need this but it doesnt work either way transforms.ToTensor()]) # Loading the data train_dataset = MNIST(os.getcwd(), train=True, download=True, transform=transform) batch_size = 10 trainloader = torch.utils.data.DataLoader(train_dataset, batch_size, shuffle=True, num_workers=1) trainloader = DeviceDataLoader(trainloader, device) # 2. The CNN model cnn_model = LeNet(lr=1e-04) to_device(cnn_model, device) # 3. Training the network # 3.1. Creating the trainer class trainer = Trainer(n_epochs=1) # 3.2. Training the model trainer.fit(cnn_model, trainloader)
排查与解决方案
1. 修改DataLoader的num_workers参数
在Windows环境下,DataLoader的多进程加载(num_workers>0)可能和CUDA存在兼容性冲突,尝试将num_workers改为0:
trainloader = torch.utils.data.DataLoader(train_dataset, batch_size, shuffle=True, num_workers=0)
2. 禁用cudnn确定性模式
RTX40系列GPU在cudnn.deterministic=True模式下可能存在已知bug,尝试注释掉该设置:
# torch.backends.cudnn.deterministic = True
3. 单独测试模型前向传播
在训练循环前添加单独的前向传播测试,确认是否是模型本身的问题:
# 生成测试输入张量 test_input = torch.randn(10, 1, 28, 28).to(device) # 无梯度模式下测试前向传播 with torch.no_grad(): output = cnn_model(test_input) print(output.shape)
如果这一步崩溃,说明模型在GPU上的初始化或运算存在问题;如果正常,再排查训练流程中的其他环节。
4. 更新PyTorch与CUDA版本
RTX4090属于较新的GPU,旧版本PyTorch对其支持可能不完善,建议更新到最新稳定版的PyTorch及对应CUDA版本。
5. 跳过Jupyter,用命令行运行脚本
将代码保存为.py文件,通过命令行执行,这样可以获取更详细的CUDA崩溃日志,帮助定位具体问题:
python lenet_train.py
内容的提问来源于stack exchange,提问作者BorisOZ
相关产品推荐
相关产品推荐

