PyTorch中LeNet5模型Swish激活函数可训练β无法更新问题排查
问题:可训练β参数的Swish激活函数未更新
我依据论文《SWISH: A Self-Gated Activation Function》,实现了带可训练β参数的Swish激活函数(替代固定β=1的nn.SiLU()),基于PyTorch 2.0和Python 3.10,用LeNet5在MNIST上做实验,但训练过程中β的值始终没有变化。
模型实现代码
class LeNet5(nn.Module): def __init__(self, beta = 1.0): super(LeNet5, self).__init__() b = torch.tensor(data = beta, dtype = torch.float32) self.beta = torch.autograd.Variable(b, requires_grad = True) self.conv1 = nn.Conv2d( in_channels = 1, out_channels = 6, kernel_size = 5, stride = 1, padding = 0, bias = False ) self.bn1 = nn.BatchNorm2d(num_features = 6) self.pool = nn.MaxPool2d(kernel_size = 2, stride = 2) self.conv2 = nn.Conv2d( in_channels = 6, out_channels = 16, kernel_size = 5, stride = 1, padding = 0, bias = False ) self.bn2 = nn.BatchNorm2d(num_features = 16) self.fc1 = nn.Linear( in_features = 256, out_features = 120, bias = True ) self.bn3 = nn.BatchNorm1d(num_features = 120) self.fc2 = nn.Linear( in_features = 120, out_features = 84, bias = True ) self.bn4 = nn.BatchNorm1d(num_features = 84) self.fc3 = nn.Linear( in_features = 84, out_features = 10, bias = True ) self.initialize_weights() def initialize_weights(self): for m in self.modules(): if isinstance(m, nn.Conv2d): nn.init.kaiming_normal_(m.weight) if m.bias is not None: nn.init.constant_(m.bias, 0) elif isinstance(m, nn.BatchNorm2d): nn.init.constant_(m.weight, 1) nn.init.constant_(m.bias, 0) elif isinstance(m, nn.Linear): nn.init.kaiming_normal_(m.weight) nn.init.constant_(m.bias, 0) def swish_fn(self, x): return x * torch.sigmoid(x * self.beta) def forward(self, x): x = self.pool(self.bn1(self.conv1(x))) x = self.swish_fn(x = x) x = self.pool(self.bn2(self.conv2(x))) x = self.swish_fn(x = x) x = x.view(-1, 256) x = self.bn3(self.fc1(x)) x = self.swish_fn(x = x) x = self.bn4(self.fc2(x)) x = self.swish_fn(x = x) x = self.fc3(x) return x
训练打印β的代码
for epoch in range(1, num_epochs + 1): train_loss, train_acc = train_one_step( model = model, train_loader = train_loader, train_dataset = train_dataset ) val_loss, val_acc = test_one_step( model = model, test_loader = test_loader, test_dataset = test_dataset ) scheduler.step() current_lr = optimizer.param_groups[0]["lr"] print(f"Epoch: {epoch}; loss = {train_loss:.4f}, acc = {train_acc:.2f}%", f" val loss = {val_loss:.4f}, val acc = {val_acc:.2f}%," f" beta = {model.beta:.6f} & LR = {current_lr:.5f}" ) train_history[epoch] = { 'train_loss': train_loss, 'val_loss': val_loss, 'train_acc': train_acc, 'val_acc': val_acc, 'lr': current_lr } if (val_acc > best_val_acc): best_val_acc = val_acc print(f"Saving model with highest val_acc = {val_acc:.2f}%\n") torch.save(model.state_dict(), "LeNet5_MNIST_best_val_acc.pth")
错误原因
β未被训练的核心问题是它没有被注册为模型的可训练参数:
- 直接创建的
torch.autograd.Variable不会自动加入模型的参数集合(model.parameters()),导致优化器无法对其进行梯度更新。 - PyTorch 2.0中
torch.autograd.Variable已被废弃,无需显式使用。
修正方案
将self.beta用nn.Parameter包裹,PyTorch会自动将其识别为模型的可训练参数,优化器就能对其进行更新。
修正后的__init__代码
def __init__(self, beta = 1.0): super(LeNet5, self).__init__() # 用nn.Parameter注册可训练参数,替代过时的Variable self.beta = nn.Parameter(torch.tensor(beta, dtype=torch.float32), requires_grad=True) # 其余层定义保持不变...
同时确保初始化优化器时,传入模型的全部参数:
optimizer = torch.optim.Adam(model.parameters(), lr=0.001)
这样训练过程中,β就会随着梯度反向传播正常更新了。
内容的提问来源于stack exchange,提问作者Arun
相关产品推荐
相关产品推荐

