You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch-Scarf运行报错:张量设备不匹配(cuda:0与cpu)

解决PyTorch-SCARF运行时的设备不匹配RuntimeError

问题情况

运行pytorch-scarf仓库的示例Notebook时触发RuntimeError,报错信息:

Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu!

仅CPU环境下运行无此问题,已确认传入损失函数的emb_anchor和emb_positive均为CUDA张量,但未确认数据加载环节的张量设备类型。

相关代码

batch_size = 128
epochs = 1000  
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

train_loader = DataLoader(train_ds, batch_size=batch_size, shuffle=True)

model = SCARF(
    input_dim=train_ds.shape[1],  
    emb_dim=16,  
    corruption_rate=0.6,
).to(device)  
optimizer = Adam(model.parameters(), lr=0.001)  
ntxent_loss = NTXent()

loss_history = []

for epoch in range(1, epochs + 1):  
    epoch_loss = train_epoch(model, ntxent_loss, train_loader, optimizer, device, epoch)  
    loss_history.append(epoch_loss)

报错堆栈

RuntimeError Traceback (most recent call last)  
Cell In [7], line 7  
  4 loss_history = []  
  6 for epoch in range(1, epochs + 1):  
----> 7 epoch_loss = train_epoch(model, ntxent_loss, train_loader, optimizer, device, epoch)  
  8 loss_history.append(epoch_loss)

File ~/pytorch-scarf/example/../example/utils.py:23, in train_epoch(model, criterion, train_loader, optimizer, device, epoch)  
  20 emb_anchor, emb_positive = model(anchor, positive)  
  22 # compute loss  
---> 23 loss = criterion(emb_anchor, emb_positive)  
  24 loss.backward()  
  26 # update model weights

File /opt/tljh/user/lib/python3.9/site-packages/torch/nn/modules/module.py:1130, in Module._call_impl(self, *input, **kwargs)  
  1126 # If we don't have any hooks, we want to skip the rest of the logic in  
  1127 # this function, and just call forward.  
  1128 if not (self._backward_hooks or self._forward_hooks or self._forward_pre_hooks or _global_backward_hooks  
  1129 or _global_forward_hooks or _global_forward_pre_hooks):  
-> 1130 return forward_call(*input, **kwargs)  
  1131 # Do not call functions when jit is used  
  1132 full_backward_hooks, non_full_backward_hooks = [], []

File ~/pytorch-scarf/example/../scarf/loss.py:39, in NTXent.forward(self, z_i, z_j)  
  37 mask = (~torch.eye(batch_size * 2, batch_size * 2, dtype=torch.bool)).float()  
  38 numerator = torch.exp(positives / self.temperature)  
---> 39 denominator = mask * torch.exp(similarity / self.temperature)  
  41 all_losses = -torch.log(numerator / torch.sum(denominator, dim=1))  
  42 loss = torch.sum(all_losses) / (2 * batch_size)

RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu!

解决方案

1. 修复损失函数中的设备不匹配问题

报错根源是NTXent损失函数中创建的mask张量默认在CPU上,而计算时的similarity张量在CUDA设备上,导致运算冲突。

修改scarf/loss.py中NTXent类的forward方法:
原代码第37行:

mask = (~torch.eye(batch_size * 2, batch_size * 2, dtype=torch.bool)).float()

修改为:

mask = (~torch.eye(batch_size * 2, batch_size * 2, dtype=torch.bool, device=z_i.device)).float()

让mask自动匹配输入张量z_i的设备。

2. 确保数据加载时张量迁移到指定设备

检查example/utils.py中的train_epoch函数,确认从DataLoader取出的批量数据已迁移到目标设备。在模型前向传播前添加:

anchor, positive = anchor.to(device), positive.to(device)

内容的提问来源于stack exchange,提问作者Ajspig

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 05:35:23