You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch构建网络时如何将神经网络深度作为参数传入?求替代方案

问题分析

你的代码里之所以不管设置多少depth,可学习参数数量都不变,核心原因是你只定义了一个隐藏层self.hidden,循环里只是重复调用这个相同的层——相当于把同一个层的参数反复用了depth次,并没有新增独立的可学习参数。

正确实现方式

下面给你两种常用的实现思路,都能让网络深度真正影响参数数量:

方法1:使用ModuleList管理多个隐藏层

ModuleList是PyTorch专门用来管理模块集合的容器,它会自动把里面的所有模块参数注册到模型中,非常适合这种动态创建多层的场景。

修改后的代码如下:

import torch
import torch.nn.functional as F

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

class Net3(torch.nn.Module):
    def __init__(self, n_feature, n_hidden, n_output, depth, init=None):
        super(Net3, self).__init__()
        self.input = torch.nn.Linear(n_feature, n_hidden).float().to(device)
        # 用ModuleList创建depth个独立的隐藏层
        self.hidden_layers = torch.nn.ModuleList([
            torch.nn.Linear(n_hidden, n_hidden).float().to(device) 
            for _ in range(depth)
        ])
        self.predict = torch.nn.Linear(n_hidden, n_output).float().to(device)
        self.depth = depth
        
        # 可选:应用初始化方法(对应你传入的init参数)
        if init is not None:
            self.input.apply(init)
            for layer in self.hidden_layers:
                layer.apply(init)
            self.predict.apply(init)

    def forward(self, x):
        x = F.relu(self.input(x))
        # 遍历每个独立的隐藏层
        for layer in self.hidden_layers:
            x = F.relu(layer(x))
        x = self.predict(x)
        return x

方法2:使用Sequential构建隐藏层序列

如果你的隐藏层结构都是统一的(比如都是Linear+ReLU),也可以用Sequential来打包这些层,代码会更简洁:

import torch
import torch.nn.functional as F

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

class Net3(torch.nn.Module):
    def __init__(self, n_feature, n_hidden, n_output, depth, init=None):
        super(Net3, self).__init__()
        layers = []
        # 输入层
        layers.append(torch.nn.Linear(n_feature, n_hidden).float().to(device))
        layers.append(torch.nn.ReLU())
        # 添加depth个隐藏层
        for _ in range(depth):
            layers.append(torch.nn.Linear(n_hidden, n_hidden).float().to(device))
            layers.append(torch.nn.ReLU())
        # 输出层
        layers.append(torch.nn.Linear(n_hidden, n_output).float().to(device))
        
        self.network = torch.nn.Sequential(*layers)
        
        # 可选:应用初始化
        if init is not None:
            self.network.apply(init)

    def forward(self, x):
        return self.network(x)
关键说明
  • 两种方法都会根据depth的数值创建对应数量的独立隐藏层,每个层都有自己的权重和偏置参数,所以参数数量会随depth增加而线性增长。
  • 原代码里的init参数之前没有被使用,上面的代码里我加上了可选的初始化逻辑,你可以根据自己的需求调整。

内容的提问来源于stack exchange,提问作者Rishith Ellath Meethal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 17:54:09