为何深度神经网络隐藏层映射的Jacobian矩阵全为零?
问题原因与解决方案
核心原因
你得到全零Jacobian矩阵的关键问题出在differentiable_hidden_mapping函数的返回值格式:该函数返回两个独立的标量张量(point[0], point[1]),而torch.autograd.functional.jacobian对这种返回元组的场景处理时,无法正确追踪完整的梯度流。虽然输出张量带有grad_fn标识梯度追踪开启,但拆分张量元素的操作干扰了自动求导系统对整个映射关系的梯度计算逻辑。
另外注意:你的differentiable_hidden_mapping没有复现forward函数中self.fc2之后的ReLU激活,这会导致隐藏层映射的中间过程与前向传播不一致,但这并非Jacobian为零的直接原因。
解决方案
修改differentiable_hidden_mapping函数,让它直接返回完整的2维张量,而非拆分元素:
def differentiable_hidden_mapping(self, point): """ @point: a tensor like torch.tensor([x, y], requires_grad=True) """ point = self.fc1(point) point = self.relu(point) point = self.fc2(point) # 若需和forward的中间过程完全一致,需添加第二个ReLU # point = self.relu(point) return point
重新执行Jacobian计算代码:
from torch.autograd.functional import jacobian point = torch.tensor([1., 1.], requires_grad=True) model.eval() print(model.differentiable_hidden_mapping(point)) print(jacobian(model.differentiable_hidden_mapping, point))
此时你会得到形状为(2, 2)的非零Jacobian矩阵,对应R²到R²映射中每个输出维度对输入维度的偏导数。
补充说明
如果确实需要保留返回拆分元素的逻辑,可以通过包装函数合并输出后再计算Jacobian:
def wrapper(point): return torch.stack(model.differentiable_hidden_mapping(point)) print(jacobian(wrapper, point))
这种方式也能正确触发梯度计算,得到非零结果。
内容的提问来源于stack exchange,提问作者Alberto
相关产品推荐
相关产品推荐

