基于向量分量定义的矩阵的PyTorch自动微分问题
问题描述
我遇到了基于向量分量定义矩阵的PyTorch自动微分相关问题,使用的代码如下:
import torch x = torch.tensor([3., 4.], requires_grad=True) A = torch.tensor([[x[0], x[1]], [x[1], x[0]]], requires_grad=True) f = torch.norm(A) f.backward()
我希望计算f对x的梯度,但调用x.grad仅得到None;若使用torch.autograd.grad(f, x)替代f.backward(),则抛出如下运行时错误:
RuntimeError: One of the differentiated Tensors appears to not have been used in the graph. Set allow_unused=True if this is the desired behavior.
问题原因
核心问题出在矩阵A的构造方式:直接用torch.tensor()包裹x的分量会打断计算图的连接。torch.tensor()会创建一个全新的张量,即使设置了requires_grad=True,这个新张量也和原x的计算图完全断开。此时f的计算图只关联到A,和x没有梯度传播路径,自然无法计算x的梯度。
解决方案
改用PyTorch原生张量操作构建矩阵,保证计算图的连续性,以下是三种可行方案:
方案1:用torch.stack构造矩阵
import torch x = torch.tensor([3., 4.], requires_grad=True) # 堆叠x的分量生成对称矩阵 A = torch.stack([torch.stack([x[0], x[1]]), torch.stack([x[1], x[0]])]) f = torch.norm(A) f.backward() print(x.grad) # 输出: tensor([0.6, 0.8])
方案2:通过赋值填充矩阵
先创建空的可微分张量,再用x的分量填充:
import torch x = torch.tensor([3., 4.], requires_grad=True) A = torch.zeros((2,2), dtype=torch.float32, requires_grad=True) A[0,0] = x[0] A[0,1] = x[1] A[1,0] = x[1] A[1,1] = x[0] f = torch.norm(A) f.backward() print(x.grad) # 输出: tensor([0.6, 0.8])
方案3:直接利用对称性计算(更高效)
由于A是对称矩阵,torch.norm(A)等价于sqrt(2*x[0]² + 2*x[1]²),可以直接对该表达式求导,无需构造矩阵:
import torch x = torch.tensor([3., 4.], requires_grad=True) f = torch.sqrt(2*(x[0]**2 + x[1]**2)) f.backward() print(x.grad) # 输出: tensor([0.6, 0.8])
梯度正确性验证
手动推导梯度:f = ||A||_F = sqrt(x₀² + x₁² + x₁² + x₀²) = sqrt(2x₀² + 2x₁²),对x₀的导数为(2x₀)/sqrt(2(x₀²+x₁²)),代入x₀=3、x₁=4,计算得6/(5√2) ≈ 0.6,与代码输出一致,验证梯度正确。
内容的提问来源于stack exchange,提问作者G. Gare

