PyTorch简易PCA类2D旋转任务中损失值未更新问题排查
嘿,这问题我太熟悉了!核心原因是你的梯度传播链路被意外切断了,导致优化器根本没办法更新模型参数,损失自然纹丝不动。咱们一步步拆解问题和修复方案:
1. 致命问题:旋转矩阵的构建切断了梯度连接
你的getrotation方法里,用torch.tensor()重新包裹了由self.theta计算出的costheta和sintheta,还手动加了requires_grad=True——这完全弄反了逻辑!
原来的costheta和sintheta是依赖self.theta的计算节点,自带梯度传播的能力,但你用torch.tensor()把它们重新打包成了全新的叶子张量,这就切断了它们和self.theta的关联。优化器只能更新self.theta,但现在梯度传不到这里,参数自然永远不变,损失也就不会下降。
修复方法:直接构建旋转矩阵,保留计算图连接
把getrotation改成这样,直接用计算得到的三角函数值构建矩阵,不要创建新的叶子张量:
def getrotation(self): sintheta = torch.sin(self.theta) costheta = torch.cos(self.theta) # 用torch.stack保留原计算图的梯度链路 return torch.stack([ torch.stack([costheta, -sintheta]), torch.stack([sintheta, costheta]) ])
这样旋转矩阵就会和self.theta保持梯度连接,反向传播时梯度就能正确传到参数上了。
2. 次要问题:变量名冲突(虽然不影响损失,但会埋坑)
你定义了covariance函数,但在训练循环里又把同名变量赋值:
covariance = (outputs[0, :] * outputs[1, :]).mean()
这会覆盖原来的函数,后续如果想调用covariance()函数就会报错。建议把变量名改成cov_loss之类的:
cov_loss = (outputs[0, :] * outputs[1, :]).mean() loss = criterion(cov_loss, torch.tensor(0, dtype=torch.double))
3. 小优化:简化输入数据处理
你原来的example2已经是torch.tensor,没必要再转一次torch.DoubleTensor,初始化时直接指定 dtype 更高效:
example2 = torch.tensor(np.random.randn(2, 33), dtype=torch.double)
4. 额外建议:调大学习率
你原来的学习率lr=0.001太小了,即使修复了梯度问题,收敛也会慢到看不见。建议调到0.1或者0.01,配合动量能更快看到损失下降:
optimizer = torch.optim.SGD(net.parameters(), lr=0.1, momentum=0.1)
修复后的完整代码
import numpy as np import torch import torch.nn as nn import torch.nn.functional as nnF class PCArot2D(nn.Module): "2D PCA rotation, expressed as a gradient-descent problem" def __init__(self): super(PCArot2D, self).__init__() # 初始化时直接指定dtype为double,避免后续转换 self.theta = nn.Parameter(torch.tensor(np.random.random() * 2 * np.pi, dtype=torch.double)) def getrotation(self): sintheta = torch.sin(self.theta) costheta = torch.cos(self.theta) # 保留梯度链路的矩阵构建方式 return torch.stack([ torch.stack([costheta, -sintheta]), torch.stack([sintheta, costheta]) ]) def forward(self, x): xmeans = torch.mean(x, dim=1, keepdim=True) rot = self.getrotation() return torch.mm(rot, x - xmeans) def covariance(y): "Calculates the covariance matrix of its input (as torch variables)" ymeans = torch.mean(y, dim=1, keepdim=True) ycentred = y - ymeans return torch.mm(ycentred, ycentred.T) / ycentred.shape[1] net = PCArot2D() # 直接初始化double类型的输入数据 example2 = torch.tensor(np.random.randn(2, 33), dtype=torch.double) # define a loss function and an optimiser criterion = nn.MSELoss() # 调大学习率,加速收敛 optimizer = torch.optim.SGD(net.parameters(), lr=0.1, momentum=0.1) # train the network num_epochs = 1000 for epoch in range(num_epochs): optimizer.zero_grad() # forward + backward + optimize outputs = net(example2) # 修复变量名冲突 cov_loss = (outputs[0, :] * outputs[1, :]).mean() loss = criterion(cov_loss, torch.tensor(0, dtype=torch.double)) loss.backward() optimizer.step() running_loss = loss.item() if ((epoch & (epoch - 1)) == 0) or epoch==(num_epochs-1): # don't print on all epochs # print statistics print('[%d] loss: %.8f' % (epoch, running_loss)) print('Finished Training')
现在运行这段代码,你应该能看到损失稳步下降,直到接近0——这才是一维PCA问题该有的样子!
内容的提问来源于stack exchange,提问作者Dan Stowell

