You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何冻结PyTorch线性层中前50行可学习权重参数?

PyTorch冻结Linear层权重前50行的可行方案

问题背景与尝试过的方法

需求:冻结nn.Linear(2048, 100)层权重的前50行(权重形状为(100, 2048))。

尝试过的方法及报错:

layer = nn.Linear(2048, 100)
# layer.weight.shape = (100, 2048)

# 尝试1:直接修改切片的requires_grad
layer.weight[:50].requires_grad = False
# 报错:RuntimeError: you can only change requires_grad flags of leaf variables. If you want to use a computed variable in a subgraph that doesn't require differentiation use var_no_grad = var.detach().

# 尝试2:修改data切片的requires_grad
layer.weight.data[:50].requires_grad = False
# 结果:layer.weight.requires_grad仍为True,未达到冻结效果

# 尝试3:用detach赋值
layer.weight[:50] = layer.weight[:50].detach()
# 报错:RuntimeError: a view of a leaf Variable that requires grad is being used in an in-place operation.

可行解决思路

方案1:反向传播时手动置零前50行梯度

这种方法不用修改参数本身的requires_grad,每次反向传播后把前50行的梯度设为0,让这部分参数不更新,实现冻结效果。

代码示例:

import torch
import torch.nn as nn

layer = nn.Linear(2048, 100)
optimizer = torch.optim.SGD(layer.parameters(), lr=0.01)

# 模拟训练流程
for _ in range(10):
    optimizer.zero_grad()
    input = torch.randn(32, 2048)
    output = layer(input)
    loss = output.sum()
    loss.backward()
    
    # 手动将前50行权重的梯度置零
    layer.weight.grad[:50] = 0.0
    
    optimizer.step()

方案2:拆分权重为可训练与不可训练参数

自定义Linear层,把权重拆成两部分:前50行(冻结,不需要梯度)和后50行(可训练),手动构建前向逻辑。

代码示例:

import torch
import torch.nn as nn

class PartialFrozenLinear(nn.Module):
    def __init__(self, in_features, out_features, freeze_rows):
        super().__init__()
        # 冻结的权重部分:不计算梯度
        self.freeze_weight = nn.Parameter(torch.randn(freeze_rows, in_features), requires_grad=False)
        # 可训练的权重部分
        self.trainable_weight = nn.Parameter(torch.randn(out_features - freeze_rows, in_features))
        self.bias = nn.Parameter(torch.randn(out_features))
        
    def forward(self, x):
        # 拼接权重
        weight = torch.cat([self.freeze_weight, self.trainable_weight], dim=0)
        return nn.functional.linear(x, weight, self.bias)

# 实例化:冻结前50行
layer = PartialFrozenLinear(2048, 100, freeze_rows=50)
# 查看可训练参数:只有trainable_weight和bias
print([name for name, param in layer.named_parameters() if param.requires_grad])
# 输出: ['trainable_weight', 'bias']

方案3:前向传播时分离前50行权重

在前向计算时,把权重的前50行用detach()分离出来(不参与梯度计算),和后50行拼接后再做线性变换。

代码示例:

import torch
import torch.nn as nn

layer = nn.Linear(2048, 100)
optimizer = torch.optim.SGD(layer.parameters(), lr=0.01)

for _ in range(10):
    optimizer.zero_grad()
    input = torch.randn(32, 2048)
    
    # 分离前50行权重,不参与梯度计算
    frozen_weight = layer.weight[:50].detach()
    trainable_weight = layer.weight[50:]
    combined_weight = torch.cat([frozen_weight, trainable_weight], dim=0)
    
    output = nn.functional.linear(input, combined_weight, layer.bias)
    loss = output.sum()
    loss.backward()
    
    optimizer.step()

内容的提问来源于stack exchange,提问作者myoh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 04:48:38