You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何对torch.cat生成的变量调用torch.autograd.grad会返回None

PyTorch autograd求导返回None的问题解析

我们已知torch.autograd.grad()会在梯度传播被阻断时返回None,以下代码的运行结果常引发疑问:

x = torch.rand(6, requires_grad=True)

y = x.pow(2).sum()
z = torch.cat([x])

grad1 = torch.autograd.grad(y, x, allow_unused=True)
grad2 = torch.autograd.grad(y, z, allow_unused=True)
      
print(f'grad1 = {grad1}, grad = {grad2}')

上述代码的输出结果为:
grad1 = (tensor([0.3705, 0.7468, 0.6102, 1.8640, 0.3518, 0.5397]),), grad = (None,)
不少开发者会认为z本质等价于x,因此grad2应该与grad1数值一致,该结果的产生原因如下:

  • x是y对应计算图的叶子节点,而z不属于y的计算图节点
  • 计算图中x是y和z共同的叶子节点,但z到y之间不存在直接梯度传播路径,因此torch.autograd.grad返回None

注意:返回的None并不代表梯度值为0

内容的提问来源于stack exchange,提问作者Tengerye

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 03:15:03