如何冻结PyTorch线性层中前50行可学习权重参数?
PyTorch冻结Linear层权重前50行的可行方案
问题背景与尝试过的方法
需求:冻结nn.Linear(2048, 100)层权重的前50行(权重形状为(100, 2048))。
尝试过的方法及报错:
layer = nn.Linear(2048, 100) # layer.weight.shape = (100, 2048) # 尝试1:直接修改切片的requires_grad layer.weight[:50].requires_grad = False # 报错:RuntimeError: you can only change requires_grad flags of leaf variables. If you want to use a computed variable in a subgraph that doesn't require differentiation use var_no_grad = var.detach(). # 尝试2:修改data切片的requires_grad layer.weight.data[:50].requires_grad = False # 结果:layer.weight.requires_grad仍为True,未达到冻结效果 # 尝试3:用detach赋值 layer.weight[:50] = layer.weight[:50].detach() # 报错:RuntimeError: a view of a leaf Variable that requires grad is being used in an in-place operation.
可行解决思路
方案1:反向传播时手动置零前50行梯度
这种方法不用修改参数本身的requires_grad,每次反向传播后把前50行的梯度设为0,让这部分参数不更新,实现冻结效果。
代码示例:
import torch import torch.nn as nn layer = nn.Linear(2048, 100) optimizer = torch.optim.SGD(layer.parameters(), lr=0.01) # 模拟训练流程 for _ in range(10): optimizer.zero_grad() input = torch.randn(32, 2048) output = layer(input) loss = output.sum() loss.backward() # 手动将前50行权重的梯度置零 layer.weight.grad[:50] = 0.0 optimizer.step()
方案2:拆分权重为可训练与不可训练参数
自定义Linear层,把权重拆成两部分:前50行(冻结,不需要梯度)和后50行(可训练),手动构建前向逻辑。
代码示例:
import torch import torch.nn as nn class PartialFrozenLinear(nn.Module): def __init__(self, in_features, out_features, freeze_rows): super().__init__() # 冻结的权重部分:不计算梯度 self.freeze_weight = nn.Parameter(torch.randn(freeze_rows, in_features), requires_grad=False) # 可训练的权重部分 self.trainable_weight = nn.Parameter(torch.randn(out_features - freeze_rows, in_features)) self.bias = nn.Parameter(torch.randn(out_features)) def forward(self, x): # 拼接权重 weight = torch.cat([self.freeze_weight, self.trainable_weight], dim=0) return nn.functional.linear(x, weight, self.bias) # 实例化:冻结前50行 layer = PartialFrozenLinear(2048, 100, freeze_rows=50) # 查看可训练参数:只有trainable_weight和bias print([name for name, param in layer.named_parameters() if param.requires_grad]) # 输出: ['trainable_weight', 'bias']
方案3:前向传播时分离前50行权重
在前向计算时,把权重的前50行用detach()分离出来(不参与梯度计算),和后50行拼接后再做线性变换。
代码示例:
import torch import torch.nn as nn layer = nn.Linear(2048, 100) optimizer = torch.optim.SGD(layer.parameters(), lr=0.01) for _ in range(10): optimizer.zero_grad() input = torch.randn(32, 2048) # 分离前50行权重,不参与梯度计算 frozen_weight = layer.weight[:50].detach() trainable_weight = layer.weight[50:] combined_weight = torch.cat([frozen_weight, trainable_weight], dim=0) output = nn.functional.linear(input, combined_weight, layer.bias) loss = output.sum() loss.backward() optimizer.step()
内容的提问来源于stack exchange,提问作者myoh
相关产品推荐
相关产品推荐

