求助:LLM中MoR模块Sigmoid输出requires_grad为False异常
梯度异常排查求助:MoR模块在LLM中梯度断裂问题
我定义了如下MoR模块:
class MoR(nn.Module): def __init__(self, work=False): super().__init__() self.sparsity = 0 self.scores = nn.Parameter(torch.tensor(-4.0)) self.f = torch.sigmoid self.threshold = self.f(self.scores) self.min_loss = float('inf') self.sparsity_best = 0 # used for training self.sparsity_loss = 0 self.recon_loss = 0 self.work = work self.num = 0 self.sparsity_avg = 0 def forward(self, inputs: torch.Tensor): if self.work: # get mask mask = self.generate_mask(inputs) if torch.sum(torch.isnan(inputs)).bool(): import pdb pdb.set_trace() self.get_sparsity(inputs) self.update_sparsity_avg() print("mask:", mask.requires_grad) print("scores:", self.scores.requires_grad) print("sparsity:", self.sparsity.requires_grad) print("threshold:", self.threshold.requires_grad) return inputs * mask else: return inputs
模块中scores的requires_grad=True,但经计算后,threshold等输出的requires_grad均为False。threshold由sigmoid(scores)计算得到,理论上完全可微分。
此异常仅在将该模块插入大语言模型(LLM)时出现,应用于如下小型模型时训练完全正常:
class LinearModel(nn.Module): def __init__(self): super(LinearModel, self).__init__() self.mor1 = MoR(work=True) self.relu1 = nn.ReLU() # Add ReLU activation layer self.linear1 = nn.Linear(1, 1) # A simple linear layer self.mor2 = MoR(work=True) self.relu2 = nn.ReLU() # Add ReLU activation layer self.linear2 = nn.Linear(1, 1) def forward(self, x): x = self.mor1(x) x = self.linear1(x) x = self.relu1(x) # Apply ReLU activation x = self.mor2(x) x = self.linear2(x) x = self.relu2(x) # Apply ReLU activation return x
已排查no_grad()上下文、计算图断裂、设备转移等常见可能原因,均未发现问题,特此求助该梯度异常的底层原因。
内容的提问来源于stack exchange,提问作者user25361577
相关产品推荐
相关产品推荐

