PyTorch中MulBackward0梯度类型不匹配RuntimeError问题求助
Hey there, let's break down this error and fix it step by step!
What's causing the error?
The error message tells us that during backpropagation (loss.backward()), a gradient tensor ended up on the CPU (torch.FloatTensor) when the rest of your model/tensors are on the GPU (torch.cuda.FloatTensor). This device mismatch breaks PyTorch's automatic differentiation pipeline.
Looking at your code snippet, the most likely culprit is the rv tensor you're initializing:
rv = torch.zeros(12 * outputs.shape[0])
By default, torch.zeros() creates a tensor on the CPU. Even though you later pass this rv to your custom ransac and predict functions (which use GPU tensors like source and target), this initial CPU tensor can introduce device mismatches in intermediate calculations, leading to the gradient error during backprop.
Step-by-step fixes
Initialize
rvon the correct device
Update thervinitialization to explicitly use your GPU device:rv = torch.zeros(12 * outputs.shape[0], device=device)This ensures
rvstarts on the GPU, matching the rest of your tensors right from the start.Check your custom functions for hidden device mismatches
Double-check yourfactor.ransacandfactor.predictfunctions. Make sure any new tensors created inside them (like withtorch.zeros,torch.rand, or numpy-to-tensor conversions) are explicitly moved to thedeviceyou're using. For example, if you have a line like:temp_tensor = torch.randn(...)Change it to:
temp_tensor = torch.randn(...).to(device)Or initialize directly on the device for efficiency:
temp_tensor = torch.randn(... , device=device)Simplify target tensor conversion
Your line for converting the target can be cleaned up to avoid unnecessary CPU operations:# Instead of this: loss = criterion(predicted, target.type(torch.FloatTensor).to(device)) # Use this (more efficient and avoids intermediate CPU tensor): loss = criterion(predicted, target.to(device, dtype=torch.float32))
Why this works
PyTorch requires all tensors involved in a computation graph to live on the same device (CPU or GPU) for backpropagation to work correctly. By ensuring every tensor—including initializations and intermediate variables in custom functions—is on the GPU, you eliminate the device mismatch that's causing the MulBackward0 gradient error.
内容的提问来源于stack exchange,提问作者Rani

