You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

复现代码时将张量移至GPU出现CUDA初始化错误求解决

问题:PyTorch数据集张量迁移至GPU后出现CUDA初始化错误

问题背景

复现PyTorch数据加载器相关代码时,将张量迁移到GPU后运行报错,但CPU环境下可正常执行。

错误信息

RuntimeError: CUDA error: initialization error
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1.
Compile with TORCH_USE_CUDA_DSA to enable device-side assertions.

出错代码

#!/usr/bin/python3

import torch
from torch.utils.data import Dataset, DataLoader
import numpy as np

if torch.cuda.is_available():
    device_ = torch.device("cuda")
    print("========================\nYou are running on GPU!\n========================")
else:
    device_ = torch.device("cpu")
    print("------------------------\nYou are running on CPU!\n------------------------")

class WineDataset(Dataset):
    
    def __init__(self):
        # data loading
        xy = np.loadtxt("./dataset/wine.csv", dtype=np.float32, delimiter=",",  skiprows=1)
        self.n_samples = xy.shape[0]

        self.x = torch.from_numpy(xy[:, 1:])
        self.x = self.x.to(device_)

        self.y = torch.from_numpy(xy[:, [0]]) # n_samples, 1
        self.y = self.y.to(device_)

        print(self.x.shape, self.y.shape)
        
    def __getitem__(self, index):
        # dataset[0]
        return self.x[index], self.y[index]

    def __len__(self):
        # len(dataset)
        return self.n_samples

dataset = WineDataset()
train_loader = DataLoader(dataset=dataset, batch_size=4, shuffle=True, num_workers=2)

dataiter = iter(train_loader)
data = next(dataiter)
features, labels = data

解决方案

问题根源是DataLoader多进程(num_workers=2)与GPU张量的冲突:Dataset的__init__方法中直接将张量移至GPU,但DataLoader的子进程无法直接访问主进程的GPU张量,从而触发CUDA初始化错误。

方案1:取数据后再迁移GPU(推荐)

先在Dataset中保留CPU张量,从DataLoader取出批量数据后再迁移到GPU:

#!/usr/bin/python3

import torch
from torch.utils.data import Dataset, DataLoader
import numpy as np

if torch.cuda.is_available():
    device_ = torch.device("cuda")
    print("========================\nYou are running on GPU!\n========================")
else:
    device_ = torch.device("cpu")
    print("------------------------\nYou are running on CPU!\n------------------------")

class WineDataset(Dataset):
    
    def __init__(self):
        # data loading
        xy = np.loadtxt("./dataset/wine.csv", dtype=np.float32, delimiter=",",  skiprows=1)
        self.n_samples = xy.shape[0]

        # 先将张量保留在CPU
        self.x = torch.from_numpy(xy[:, 1:])
        self.y = torch.from_numpy(xy[:, [0]]) # n_samples, 1

        print(self.x.shape, self.y.shape)
        
    def __getitem__(self, index):
        return self.x[index], self.y[index]

    def __len__(self):
        return self.n_samples

dataset = WineDataset()
train_loader = DataLoader(dataset=dataset, batch_size=4, shuffle=True, num_workers=2)

dataiter = iter(train_loader)
data = next(dataiter)
features, labels = data
# 取出批量数据后迁移至GPU
features = features.to(device_)
labels = labels.to(device_)

方案2:在__getitem__内迁移GPU

每次获取单条数据时,将其迁移到GPU:

# 修改WineDataset类的__getitem__方法
def __getitem__(self, index):
    x_item = self.x[index].to(device_)
    y_item = self.y[index].to(device_)
    return x_item, y_item

备选方案(不推荐)

如果不需要多进程加速,可以设置num_workers=0禁用多进程,但会降低数据加载效率。

内容的提问来源于stack exchange,提问作者Bilal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 19:08:10