You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch模型与张量均在GPU却触发设备不匹配RuntimeError

问题:PyTorch ConvNet推理触发设备不匹配RuntimeError

执行代码outputs = model(X_batch)时出现以下错误:

RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cpu and cuda:0! (when checking argument for argument mat1 in method wrapper_CUDA_addmm)

已确认模型和输入张量均已移至GPU:

print(args.device) # 返回 'cuda'
print(torch.cuda.is_available()) # 返回 True

model = CNN_MLP(args)
model.to(args.device)

def inference(model, test_dataloader, device='cpu'):
    model.eval()

    metrics = []
    for _, (X_batch, y_batch) in enumerate(tqdm(test_dataloader)):
        X_batch = X_batch.to(args.device)

        print(next(model.parameters()).is_cuda) # 返回 True
        print(X_batch.is_cuda) # 返回 True

        outputs = model(X_batch) # 此处触发错误
        metrics.append(calc_metrics(outputs, y_batch))

    return aggregate(metrics)

zero_shot_metrics = inference(model, test_dataloader, device=args.device)

问题原因

你的CNN_MLP类中,fc_layers是用普通Python列表存储的,没有通过torch.nn.Sequential包装。PyTorch的model.to(device)只会移动注册为模型子模块的参数,列表中的Linear、Dropout等层并未被注册,因此仍然留在CPU上。当forward过程中GPU上的张量传入这些CPU层时,就会触发设备不匹配错误。

修复方案

将fc_layers改为用nn.Sequential包装,让PyTorch自动管理这些子模块:

修改CNN_MLP的__init__部分

# 替换原fc_layers构建代码
self.hidden_layer_sizes.insert(0, self._linear_layer_in_size())
fc_layers_list = []
for i in range(len(self.hidden_layer_sizes) - 1):
    fc_layers_list.append(torch.nn.Linear(self.hidden_layer_sizes[i], self.hidden_layer_sizes[i+1]))
    fc_layers_list.append(torch.nn.ReLU())
    if self.dropout_rate and i != len(self.hidden_layer_sizes) - 2:
        fc_layers_list.append(torch.nn.Dropout(self.dropout_rate))
# 用nn.Sigmoid()模块替代torch.sigmoid函数,确保能被设备迁移
fc_layers_list.append(torch.nn.Sigmoid())
self.fc_layers = torch.nn.Sequential(*fc_layers_list)

修改forward方法中的层调用逻辑

# 替换原循环遍历代码
x = self.fc_layers(x)

额外验证

修复后可添加以下代码,确认所有模型参数都已迁移到GPU:

for name, param in model.named_parameters():
    print(f"{name}: {param.device}")

所有参数的设备应显示为cuda:0。

内容的提问来源于stack exchange,提问作者Vedang Waradpande

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 13:55:33