You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

仅含GRU层的模型为何返回零梯度?梯度行为探究

GRU模型梯度行为疑问解答

模型定义

第一个模型(含GRU与全连接层)

import torch
import torch.nn as nn
import torchinfo

class MyModel(nn.Module):
    def __init__(self):
        super(MyModel, self).__init__()        
        self.Identity =  nn.Identity ()
        self.GRU      = nn.GRU(input_size=3, hidden_size=32, num_layers=2, batch_first=True)
        self.fc       = nn.Linear(32, 5)
        
    def forward(self, input_series):
                
        self.Identity(input_series)
        
        output, h = self.GRU(input_series)                
        output    = output[:,  -1, :]       # 获取最后一个状态                        
        output    = self.fc(output) 
        output    = output.view(-1, 5, 1)   # 重组输出                
                        
        return output

第二个模型(仅GRU层)

class SecondModel(nn.Module):
    def __init__(self):
        super(SecondModel, self).__init__()        
        self.GRU      = nn.GRU(input_size=3, hidden_size=32, num_layers=2, batch_first=True)        
        
    def forward(self, input_series):
                
        output, h = self.GRU(input_series)                                        
        return output

梯度测试代码及结果

第一个模型测试

model = MyModel()
x     = torch.rand([2, 10, 3])
y     = model(x)
y.retain_grad()  
y[:, -1].sum().backward()
print(torch.allclose(y.grad[:, :-1], torch.tensor(0.)))  # 输出True,前N-1个输出梯度为零

第二个模型测试

model = SecondModel()
x     = torch.rand([2, 10, 3])
y     = model(x)
y.retain_grad()  
y[:, -1].sum().backward()
print(torch.allclose(y.grad[:, :-1], torch.tensor(0.)))  # 输出True,前N-1个输出梯度为零

疑问解答

1. 忽略的核心要点

你混淆了GRU输出张量的梯度与GRU模型参数的梯度的区别:

  • 当前代码计算的是模型输出张量y的梯度(通过y.retain_grad()),而非GRU层权重、偏置等参数的梯度。
  • 对于仅GRU的模型,损失只用到了最后一个时间步的输出,前N-1个时间步的输出完全没参与损失计算,所以这些前序输出的梯度自然为零。
  • Stack Overflow提到的“非零梯度”,指的是GRU层参数的梯度,你可以通过打印model.GRU.weight_hh_l0.grad验证,这类参数的梯度是不为零的。

另外第一个模型中,你直接丢弃了GRU前序时间步的输出,只取最后一步传入全连接层,前序输出未参与损失计算,梯度为零完全合理。

2. 零梯度/非零梯度的场景

  • 输出张量的零梯度场景:
    • 部分输出未参与损失函数计算,比如仅用最后一个时间步输出求和,前序时间步输出梯度为零。
    • 张量经过detach()操作,后续计算与该张量无关,梯度为零。
  • 输出张量的非零梯度场景:
    • 输出张量的所有部分都参与损失计算,比如对整个y求和后反向传播,此时y全位置梯度非零。
  • 模型参数的零梯度场景:
    • 参数对应的计算路径未参与损失正向传播,比如第一个模型的Identity层(无参数),或某层输出被完全丢弃。
    • 梯度被手动清零(如optimizer.zero_grad())。
  • 模型参数的非零梯度场景:
    • 参数参与了损失的正向传播路径,且反向传播未被阻断,比如第二个模型中GRU的所有参数都参与了最后一个时间步输出的计算,梯度非零。

内容的提问来源于stack exchange,提问作者user3668129

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 07:01:08