PyTorch中两个模型的autograd.grad()计算精度差异问题
我用PyTorch实现了两个神经网络,二者结构与权重初始化完全一致,唯一区别是其中一个模型在forward方法末尾添加了自定义插值层。
自定义插值逻辑及带插值层的模型代码如下:
def custom_output(o, x, y): minA = [-0.5, 1] minB = [0.5, 1] XA = (1/2 -1/2 * torch.tanh(1000*(((x - minA[0])**2 + (y - minA[1])**2) - (0.1 + 0.02)**2))) XB = (1/2 -1/2 * torch.tanh(1000*(((x - minB[0])**2 + (y - minB[1])**2) - (0.1 + 0.02)**2))) q = (1-XA)*((1-XB)* o + (XB)) return q class CustomLayer(nn.Module): def __init__(self): super(CustomLayer, self).__init__() def forward(self, input, x, y): # Apply custom function using input, x, and y output = custom_output(input, x, y) return output class InterpolatedNN(nn.Module): def __init__(self, input_size): #number of initial nodes super(InterpolatedNN, self).__init__() #calls the initiation function of the class self.hidden1 = nn.Linear(input_size, 10) # 2 input features (x and y coordinates), 10 hidden units self.hidden2 = nn.Linear(10, 20) self.output = nn.Linear(20, 1) self.interpolation = CustomLayer() def forward(self, x, y): coordinates = torch.cat((x, y), dim=1) out = torch.tanh(self.hidden1(coordinates)) out = torch.tanh(self.hidden2(out)) out = torch.sigmoid(self.output(out)) out = self.interpolation(out, x, y) return out
基础模型代码如下:
class TwoLayersNN(nn.Module): def __init__(self, input_size): #number of initial nodes super(TwoLayersNN, self).__init__() #calls the initiation function of the class self.hidden1 = nn.Linear(input_size, 10) # 2 input features (x and y coordinates), 10 hidden units self.hidden2 = nn.Linear(10, 20) self.output = nn.Linear(20, 1) def forward(self, x, y): coordinates = torch.cat((x, y), dim=1) out = torch.tanh(self.hidden1(coordinates)) out = torch.tanh(self.hidden2(out)) out = torch.sigmoid(self.output(out)) return out
执行以下代码验证两个模型的输出梯度:
x = torch.rand(10, 1).requires_grad_() y = torch.rand(10, 1).requires_grad_() torch.manual_seed(1) m_1 = InterpolatedNN(2) out = m_1(x, y) o = custom_output(out, x, y) grad_x, grad_y = torch.autograd.grad(out.sum(), (x, y), create_graph=True) print(grad_x) torch.manual_seed(1) m_2 = TwoLayersNN(2) out = m_2(x, y) grad_x, grad_y = torch.autograd.grad(out.sum(), (x, y), create_graph=True) print(grad_x)
得到的梯度输出存在明显精度差异:
tensor([[-8.2997e-03], [-5.3281e-02], [-6.3462e-03], [-1.1440e-02], [-1.1525e-02], [-1.4034e-02], [ 8.6983e-06], [-1.1623e-02], [-9.2993e-03], [-1.0333e-02]], grad_fn=<AddBackward0>) tensor([[-0.0083], [-0.0083], [-0.0063], [-0.0114], [-0.0115], [-0.0140], [-0.0089], [-0.0116], [-0.0093], [-0.0103]], grad_fn=<SliceBackward0>)
请问为何两个模型的autograd.grad()计算结果会存在精度差异?
梯度差异的核心原因是自定义插值层里的tanh(1000*...)操作引入了严重的数值不稳定,具体细节如下:
tanh函数的饱和特性
tanh函数在输入绝对值很大时,输出会趋近于±1,此时它的梯度(1 - tanh(x)^2)会趋近于0。你的代码里用了1000作为放大系数,这意味着只要((x - minA[0])**2 + (y - minA[1])**2) - (0.1 + 0.02)**2的结果稍微偏离0,乘以1000后就会让tanh的输入进入绝对值极大的区域,导致梯度几乎消失。浮点精度的微小波动被放大
即使输入值非常接近tanh的非饱和区,1000倍的放大系数也会把浮点计算中的微小精度误差放大,导致tanh的输出出现异常跳变,进而让反向传播时的梯度计算出现偏差。比如你看到的第二个梯度值-5.3281e-02和基础模型的-0.0083差异巨大,就是因为对应的(x,y)刚好落在了tanh饱和区的边缘,数值误差被放大后导致梯度异常。自定义层对梯度传导的影响
基础模型的梯度直接来自sigmoid和tanh的反向传播,数值稳定;而带插值层的模型,梯度需要经过custom_output中的一系列运算,尤其是经过tanh(1000*...)这个数值不稳定的环节,梯度的计算精度被严重破坏,甚至出现梯度消失或异常值(比如第七个梯度值8.6983e-06几乎为0,和基础模型的-0.0089完全不符)。
如果要解决这个问题,可以尝试降低tanh的输入放大系数(比如把1000改成100或10),或者用更平滑的函数替代tanh实现类似的门控逻辑,避免进入饱和区导致数值不稳定。
内容的提问来源于stack exchange,提问作者eliss

