PyTorch RuntimeError:张量设备不一致问题求助及张量查询咨询
问题1:RuntimeError 原因及解决方法
错误根源
你的代码存在两处设备不匹配问题:
- 输入张量
src未转移到CUDA:你通过model.to(device)将模型参数移到了CUDA设备,但输入的src默认创建在CPU上,传入模型后forward中的x为CPU张量,与CUDA上的模型参数计算时触发设备冲突。 forward中src与tgt设备不一致:虽然tgt通过.to(device)转到了CUDA,但x(即src)仍在CPU,而Transformer要求输入的src和tgt必须处于同一设备,这也会引发错误。
修复后的代码
import torch.nn as nn import torch class MyModel(nn.Module): def __init__(self): super().__init__() self.lin = nn.Linear(5,1) self.trans = nn.Transformer(nhead=1, num_encoder_layers=1, d_model=5) def forward(self, x): # 基于x的设备创建tgt,避免硬编码设备 tgt = torch.rand(4, 5).to(x.device) y = self.trans(x, tgt) out = self.lin(y) return out device = torch.device('cuda:0' if torch.cuda.is_available() else 'cpu') model = MyModel() model.to(device) model.eval() # 将输入src转移到指定设备 src = torch.rand(4,5).to(device) out = model(src) print(out)
问题2:列出某设备上的所有张量
PyTorch没有内置命令直接列出指定设备上的所有张量,需通过内存对象遍历实现。可以借助gc模块筛选出目标设备的张量:
import gc import torch def list_tensors_on_device(target_device): target_tensors = [] for obj in gc.get_objects(): try: if isinstance(obj, torch.Tensor) and not obj.is_meta and obj.device == target_device: target_tensors.append(obj) except Exception: # 跳过无法判断的对象 continue return target_tensors # 示例:列出cuda:0上的所有张量 device = torch.device('cuda:0') tensors_on_cuda = list_tensors_on_device(device) print(f"{device}上的张量数量:{len(tensors_on_cuda)}")
注意:该方法会列出所有存活的张量对象,包括模型参数、输入输出、临时计算张量等,结果可能包含大量框架内部使用的张量,需自行按需筛选。
内容的提问来源于stack exchange,提问作者JobHunter69
相关产品推荐
相关产品推荐

