PyTorch RuntimeError求助:张量设备不匹配(cuda:0与cpu)问题排查
解决PyTorch中
Expected all tensors to be on the same device错误 你的问题根源在于模型里的隐藏层没有被PyTorch正确注册为子模块,导致调用model.to(device)时,这些层的参数还留在CPU上,而输入张量已经传到了GPU,两者设备不匹配就抛出了这个错误。
具体看你的MD5EncryptionEncoder类:你用了普通的Python列表self.hidden_layers来存储Linear层,但PyTorch只会自动管理通过self.layer_name = nn.Layer()这种方式注册的子模块,普通列表里的层不会被框架识别,所以当你把模型移到GPU时,这些隐藏层的权重和偏置还留在CPU。当你在forward里用这些层处理GPU上的输入张量时,就会触发设备不匹配的错误。
修复方案
把self.hidden_layers改成nn.ModuleList,这样PyTorch就能识别并管理这些子模块,调用model.to(device)时会自动把所有层的参数转移到目标设备。
修改后的MD5EncryptedDataEncoder.py代码:
import torch import torch.nn.functional as F from torch import nn # define the network class class MD5EncryptionEncoder(nn.Module): m_save_path = "data/MD5EncryptionEncoder.model" def __init__(self): # call constructor from superclass super().__init__() self.input_to_hidden_1 = nn.Linear(128, 10240) # 把普通列表改成ModuleList,让PyTorch管理这些子模块 self.hidden_layers = nn.ModuleList([ nn.Linear(10240, 10240), nn.Linear(10240, 10240) ]) self.last_hidden_to_output = nn.Linear(10240, 128) self.hidden_layers_activation_function = torch.nn.LeakyReLU(0.1) def forward(self, x): # define forward pass. # Here, 'x' represent the output of a network defined in __init__ x = self.hidden_layers_activation_function(self.input_to_hidden_1(x)) x = self.hidden_layers_activation_function(self.hidden_layers[0](x)) x = self.hidden_layers_activation_function(self.hidden_layers[1](x)) x = torch.sigmoid(self.last_hidden_to_output(x)) return x def EncryptedTensorToStateTensor(self,x): x = self.hidden_layers_activation_function(self.input_to_hidden_1(x)) x = self.hidden_layers_activation_function(self.hidden_layers[0](x)) return x def StateTensorToEncryptedTensor(self,x): x = self.hidden_layers_activation_function(self.hidden_layers[1](x)) x = torch.sigmoid(self.last_hidden_to_output(x)) return x
额外注意事项
- 如果你之前已经保存了模型,加载的时候可能会因为模型结构变化(从列表变成ModuleList)出现问题,这时候有两个选择:
- 重新初始化模型,从头开始训练(推荐,因为之前的模型参数没有正确转移到GPU,训练效果可能也有问题)
- 如果要加载旧模型,可以在加载后手动把hidden_layers里的层转移到设备:
if model_file.is_file(): model = torch.load(model.m_save_path) print("Loaded previously saved model from ["+model.m_save_path+"]") # 手动把旧模型里的hidden_layers转移到device for layer in model.hidden_layers: layer.to(device)
验证修复
修改后,你可以在训练代码里加一句打印,确认所有子模块都在GPU上:
print("All submodule devices:") for name, module in model.named_modules(): if hasattr(module, 'weight'): print(f"{name}: {module.weight.device}")
这样就能看到所有Linear层的参数都在cuda:0上了,不会再出现设备不匹配的错误。
内容的提问来源于stack exchange,提问作者Charles-Ugo Brouillard
相关产品推荐
相关产品推荐

