PyTorch nn.Linear层输入权重正常却输出NaN的Bug求助
Hey everyone, I'm stuck on a super odd PyTorch bug and could use some fresh eyes. Let me break down what's happening:
I've got a neural network with a fully connected layer net.fc_h1 (using nn.Linear). During training, I noticed this layer was spitting out NaNs right before I apply the tanh activation. To catch this, I added a check in my forward pass:
def forward(self, obs): z1 = self.fc_h1(obs) if np.isnan(np.sum(z1.data.numpy())): pdb.set_trace() h1 = F.tanh(z1) # Remaining forward pass logic...
Sure enough, the debugger triggers when a NaN is detected. But here's the kicker—when I run the exact same layer computation again inside pdb, it returns a perfectly valid value:
(Pdb) z1.sum()
Variable containing:
nan
[torch.FloatTensor of size 1](Pdb) self.fc_h1(obs).sum()
Variable containing:
771.5120
[torch.FloatTensor of size 1]
I've been scratching my head over this. Some things I've already checked or considered:
- Non-determinism: I disabled cuDNN's non-deterministic operations, but the bug still pops up.
- Parameter integrity: I inspected
net.fc_h1.weightandnet.fc_h1.biasright when the NaN is caught—they don't have any NaNs or infinities, and look identical to when I recompute the layer output. - Floating-point edge cases: Could this be a rare precision issue that only hits under specific training conditions? The input
obsalso doesn't have any NaNs when I check it in pdb. - Parallel training quirks: I'm using single-GPU training right now, so race conditions from multi-device setups shouldn't be the issue.
Has anyone encountered something like this before? Any ideas on what could be causing this temporary NaN that vanishes on recomputation?
内容的提问来源于stack exchange,提问作者Z. Liu

