You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch autograd无指数化变量时返回None或报错的原因咨询

PyTorch autograd计算二阶偏导的异常行为解析

案例1:返回x**1时输出tensor([0., 0.])

代码:

def f(x,y):
   return x**1 
x = torch.tensor([3.0,2.0], requires_grad=True)
y = torch.tensor([2.0,3.0], requires_grad=True)
z = f(x,y)
# differentiate z with respect to x twice
for _ in range(2):
    dx = torch.autograd.grad(z, x, grad_outputs=torch.ones_like(x), create_graph=True, allow_unused=True)[0]
    z = dx
print(z)

原因:
虽然x**1数学上等价于x,但PyTorch会为这个幂运算创建PowBackward0类型的梯度函数(grad_fn)。第一次对z求导得到全1的张量,且这个张量保留了计算图(因为create_graph=True);第二次对这个全1张量求导,结果就是全0张量,所以最终输出tensor([0., 0.])。


案例2:返回x*y时输出None

代码:

def f(x,y):
    return x*y 
x = torch.tensor([3.0,2.0], requires_grad=True)
y = torch.tensor([2.0,3.0], requires_grad=True)
z = f(x,y)
# differentiate z with respect to x twice
for _ in range(2):
    dx = torch.autograd.grad(z, x, grad_outputs=torch.ones_like(x), create_graph=True, allow_unused=True)[0]
    z = dx
print(z)

原因:
第一次求导时,x*y对x的偏导是y,所以dx等于y;第二次求导是对y求关于x的偏导,而y和x是独立的张量,不存在梯度关系。加上参数allow_unused=True,PyTorch会对这种无梯度的情况返回None,所以最终输出None。


案例3:直接返回x时抛出RuntimeError

代码:

def f(x,y):
   return x 
x = torch.tensor([3.0,2.0], requires_grad=True)
y = torch.tensor([2.0,3.0], requires_grad=True)
z = f(x,y)
# differentiate z with respect to x twice
for _ in range(2):
    dx = torch.autograd.grad(z, x, grad_outputs=torch.ones_like(x), create_graph=True, allow_unused=True)[0]
    z = dx
print(z)

错误信息:

RuntimeError                              Traceback (most recent call last)
Cell In[59], line 9
  7 # differentiate z with respect to x twice
  8 for _ in range(2):
----> 9     dx = torch.autograd.grad(z, x, grad_outputs=torch.ones_like(x), create_graph=True, allow_unused=True)[0]
 10     z = dx
 11 print(z)

 File ~/miniconda3/envs/randomvenv/lib/python3.8/site-packages/torch/autograd/__init__.py:394, in grad(outputs, inputs, grad_outputs, retain_graph, create_graph, only_inputs, allow_unused, is_grads_batched, materialize_grads)
390     result = _vmap_internals._vmap(vjp, 0, 0, allow_none_pass_through=True)(
391         grad_outputs_
392     )
393 else:
--> 394     result = Variable._execution_engine.run_backward(  # Calls into the C++ engine to run the backward pass
395         t_outputs,
396         grad_outputs_,
397         retain_graph,
398         create_graph,
399         t_inputs,
400         allow_unused,
401         accumulate_grad=False,
402     )  # Calls into the C++ engine to run the backward pass
403 if materialize_grads:
404     result = tuple(
405         output
406         if output is not None
407         else torch.zeros_like(input, requires_grad=True)
408         for (output, input) in zip(result, t_inputs)
409     )

RuntimeError: element 0 of tensors does not require grad and does not have a grad_fn

原因:
第一次求导时,z就是x本身,对x求导得到的是全1的张量,但这个张量是直接由梯度计算生成的常数张量,没有关联的梯度函数(grad_fn),且requires_grad默认为False。第二次循环时,用这个张量作为z去求导,create_graph=True要求被求导的张量必须可导(要么requires_grad=True,要么有grad_fn),但这个张量两者都不满足,所以抛出RuntimeError。


内容的提问来源于stack exchange,提问作者Toonia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 11:57:30