复现代码时将张量移至GPU出现CUDA初始化错误求解决
问题:PyTorch数据集张量迁移至GPU后出现CUDA初始化错误
问题背景
复现PyTorch数据加载器相关代码时,将张量迁移到GPU后运行报错,但CPU环境下可正常执行。
错误信息
RuntimeError: CUDA error: initialization error
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1.
Compile withTORCH_USE_CUDA_DSAto enable device-side assertions.
出错代码
#!/usr/bin/python3 import torch from torch.utils.data import Dataset, DataLoader import numpy as np if torch.cuda.is_available(): device_ = torch.device("cuda") print("========================\nYou are running on GPU!\n========================") else: device_ = torch.device("cpu") print("------------------------\nYou are running on CPU!\n------------------------") class WineDataset(Dataset): def __init__(self): # data loading xy = np.loadtxt("./dataset/wine.csv", dtype=np.float32, delimiter=",", skiprows=1) self.n_samples = xy.shape[0] self.x = torch.from_numpy(xy[:, 1:]) self.x = self.x.to(device_) self.y = torch.from_numpy(xy[:, [0]]) # n_samples, 1 self.y = self.y.to(device_) print(self.x.shape, self.y.shape) def __getitem__(self, index): # dataset[0] return self.x[index], self.y[index] def __len__(self): # len(dataset) return self.n_samples dataset = WineDataset() train_loader = DataLoader(dataset=dataset, batch_size=4, shuffle=True, num_workers=2) dataiter = iter(train_loader) data = next(dataiter) features, labels = data
解决方案
问题根源是DataLoader多进程(num_workers=2)与GPU张量的冲突:Dataset的__init__方法中直接将张量移至GPU,但DataLoader的子进程无法直接访问主进程的GPU张量,从而触发CUDA初始化错误。
方案1:取数据后再迁移GPU(推荐)
先在Dataset中保留CPU张量,从DataLoader取出批量数据后再迁移到GPU:
#!/usr/bin/python3 import torch from torch.utils.data import Dataset, DataLoader import numpy as np if torch.cuda.is_available(): device_ = torch.device("cuda") print("========================\nYou are running on GPU!\n========================") else: device_ = torch.device("cpu") print("------------------------\nYou are running on CPU!\n------------------------") class WineDataset(Dataset): def __init__(self): # data loading xy = np.loadtxt("./dataset/wine.csv", dtype=np.float32, delimiter=",", skiprows=1) self.n_samples = xy.shape[0] # 先将张量保留在CPU self.x = torch.from_numpy(xy[:, 1:]) self.y = torch.from_numpy(xy[:, [0]]) # n_samples, 1 print(self.x.shape, self.y.shape) def __getitem__(self, index): return self.x[index], self.y[index] def __len__(self): return self.n_samples dataset = WineDataset() train_loader = DataLoader(dataset=dataset, batch_size=4, shuffle=True, num_workers=2) dataiter = iter(train_loader) data = next(dataiter) features, labels = data # 取出批量数据后迁移至GPU features = features.to(device_) labels = labels.to(device_)
方案2:在__getitem__内迁移GPU
每次获取单条数据时,将其迁移到GPU:
# 修改WineDataset类的__getitem__方法 def __getitem__(self, index): x_item = self.x[index].to(device_) y_item = self.y[index].to(device_) return x_item, y_item
备选方案(不推荐)
如果不需要多进程加速,可以设置num_workers=0禁用多进程,但会降低数据加载效率。
内容的提问来源于stack exchange,提问作者Bilal
相关产品推荐
相关产品推荐

