PyTorch报错:‘One of the differentiated Tensors appears to not have been used in the graph’求助
问题分析与解决方案
报错原因
你遇到的问题核心是:第一次计算的output对x的梯度本身与x无关,导致后续y_hat(梯度之和)和x之间没有计算图依赖。
你的模型是纯线性结构,输出可展开为:output = x @ weight1.T @ weight2.T。对x求导后得到的梯度是weight1 @ weight2——这是一个完全由模型参数决定的常数张量,和输入x没有任何关联。当你把这个梯度求和得到y_hat时,y_hat已经是与x无关的常数,此时计算y_hat对x的梯度,PyTorch会判定x没有参与y_hat的计算图构建,因此抛出报错。
修正方案
情况1:目标是计算输出对x的二阶导数(Hessian矩阵的迹)
对于纯线性模型,二阶导数本身为0,但可以用torch.autograd.functional.hessian直接实现需求:
import torch import torch.nn as nn import torch.nn.functional as F from torch.autograd.functional import hessian class Model(nn.Module): def __init__(self,): super(Model, self).__init__() self.weight1 = torch.nn.Parameter(torch.tensor([[.2,.5,.9],[1.0,.3,.5],[.3,.2,.7]])) self.weight2 = torch.nn.Parameter(torch.tensor([2.0,1.0,.4])) def forward(self, x): out = F.linear(x, self.weight1.T) out = F.linear(out, self.weight2.T) return out model = Model() x = torch.tensor([[0.1,0.7,0.2]], requires_grad=True) # 定义辅助函数,返回输出的标量和(hessian要求输出为标量) def func(x): return model(x).sum() # 计算Hessian矩阵并求迹(即所有二阶导数的和) hess = hessian(func, x) y_hat_hess_trace = hess.sum() # 对x求导,线性模型下结果为0 grad_yhat_x = torch.autograd.grad(y_hat_hess_trace, x) print(grad_yhat_x) # 输出 (tensor([[0., 0., 0.]]),)
情况2:模型后续会加入非线性层(梯度将依赖x)
如果你的模型后续会引入非线性操作(比如ReLU),此时输出对x的梯度会依赖x本身,原代码逻辑可以调整后正常运行:
import torch import torch.nn as nn import torch.nn.functional as F class Model(nn.Module): def __init__(self,): super(Model, self).__init__() self.weight1 = torch.nn.Parameter(torch.tensor([[.2,.5,.9],[1.0,.3,.5],[.3,.2,.7]])) self.weight2 = torch.nn.Parameter(torch.tensor([2.0,1.0,.4])) def forward(self, x): out = F.linear(x, self.weight1.T) out = F.relu(out) # 加入非线性层,让梯度依赖x out = F.linear(out, self.weight2.T) return out model = Model() x = torch.tensor([[0.1,0.7,0.2]], requires_grad=True) output = model(x) # 计算output对x的梯度并保留计算图 grad_output_x, = torch.autograd.grad(output.sum(), x, create_graph=True) y_hat = grad_output_x.sum() # 此时y_hat依赖x,可正常求导 grad_yhat_x, = torch.autograd.grad(y_hat, x) print(grad_yhat_x)
关键总结
- 纯线性模型中,输出对输入的梯度是常数,与输入无关,因此无法基于该梯度构建关于输入的计算图;
- 若需求是二阶导数,直接使用
hessian工具更简洁; - 只有当模型包含非线性层时,输入的梯度才会依赖输入本身,原代码逻辑才能正常运行。
内容的提问来源于stack exchange,提问作者rene smith
相关产品推荐
相关产品推荐

