PyTorch autograd无指数化变量时返回None或报错的原因咨询
PyTorch autograd计算二阶偏导的异常行为解析
案例1:返回x**1时输出tensor([0., 0.])
代码:
def f(x,y): return x**1 x = torch.tensor([3.0,2.0], requires_grad=True) y = torch.tensor([2.0,3.0], requires_grad=True) z = f(x,y) # differentiate z with respect to x twice for _ in range(2): dx = torch.autograd.grad(z, x, grad_outputs=torch.ones_like(x), create_graph=True, allow_unused=True)[0] z = dx print(z)
原因:
虽然x**1数学上等价于x,但PyTorch会为这个幂运算创建PowBackward0类型的梯度函数(grad_fn)。第一次对z求导得到全1的张量,且这个张量保留了计算图(因为create_graph=True);第二次对这个全1张量求导,结果就是全0张量,所以最终输出tensor([0., 0.])。
案例2:返回x*y时输出None
代码:
def f(x,y): return x*y x = torch.tensor([3.0,2.0], requires_grad=True) y = torch.tensor([2.0,3.0], requires_grad=True) z = f(x,y) # differentiate z with respect to x twice for _ in range(2): dx = torch.autograd.grad(z, x, grad_outputs=torch.ones_like(x), create_graph=True, allow_unused=True)[0] z = dx print(z)
原因:
第一次求导时,x*y对x的偏导是y,所以dx等于y;第二次求导是对y求关于x的偏导,而y和x是独立的张量,不存在梯度关系。加上参数allow_unused=True,PyTorch会对这种无梯度的情况返回None,所以最终输出None。
案例3:直接返回x时抛出RuntimeError
代码:
def f(x,y): return x x = torch.tensor([3.0,2.0], requires_grad=True) y = torch.tensor([2.0,3.0], requires_grad=True) z = f(x,y) # differentiate z with respect to x twice for _ in range(2): dx = torch.autograd.grad(z, x, grad_outputs=torch.ones_like(x), create_graph=True, allow_unused=True)[0] z = dx print(z)
错误信息:
RuntimeError Traceback (most recent call last) Cell In[59], line 9 7 # differentiate z with respect to x twice 8 for _ in range(2): ----> 9 dx = torch.autograd.grad(z, x, grad_outputs=torch.ones_like(x), create_graph=True, allow_unused=True)[0] 10 z = dx 11 print(z) File ~/miniconda3/envs/randomvenv/lib/python3.8/site-packages/torch/autograd/__init__.py:394, in grad(outputs, inputs, grad_outputs, retain_graph, create_graph, only_inputs, allow_unused, is_grads_batched, materialize_grads) 390 result = _vmap_internals._vmap(vjp, 0, 0, allow_none_pass_through=True)( 391 grad_outputs_ 392 ) 393 else: --> 394 result = Variable._execution_engine.run_backward( # Calls into the C++ engine to run the backward pass 395 t_outputs, 396 grad_outputs_, 397 retain_graph, 398 create_graph, 399 t_inputs, 400 allow_unused, 401 accumulate_grad=False, 402 ) # Calls into the C++ engine to run the backward pass 403 if materialize_grads: 404 result = tuple( 405 output 406 if output is not None 407 else torch.zeros_like(input, requires_grad=True) 408 for (output, input) in zip(result, t_inputs) 409 ) RuntimeError: element 0 of tensors does not require grad and does not have a grad_fn
原因:
第一次求导时,z就是x本身,对x求导得到的是全1的张量,但这个张量是直接由梯度计算生成的常数张量,没有关联的梯度函数(grad_fn),且requires_grad默认为False。第二次循环时,用这个张量作为z去求导,create_graph=True要求被求导的张量必须可导(要么requires_grad=True,要么有grad_fn),但这个张量两者都不满足,所以抛出RuntimeError。
内容的提问来源于stack exchange,提问作者Toonia
相关产品推荐
相关产品推荐

