You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

两类深度神经网络初始化代码为何表现不同?

深度神经网络初始化方法的形状不匹配问题

我在实现简单深度神经网络时,遇到一个奇怪的问题:两种逻辑上看似等价的__init__权重初始化方式,实际运行时表现完全不同。第二种版本正常工作,但第一种在执行前向传播时会触发形状不匹配错误。

第一种初始化版本的完整类

import numpy as np

class DeepNeuralNetwork():
   def __init__(self, nx, layers):
        """
        @nx is the number of input features
        @layers is a list representing the number
        of nodes in each layer of the network
        """
        self.__L = len(layers)
        self.__cache = {}
        self.__weights = {}
        layers.insert(0, nx)

        for i in range(1, len(layers)):
            self.__weights["W{}".format(i)] = np.random.randn(
                layers[i], layers[i - 1]) * np.sqrt(2 / layers[i - 1])
            self.__weights["b{}".format(i)] = np.zeros((layers[i], 1))


    def forward_prop(self, X):
        """ Calculates forward propagation """
        self.__cache["A0"] = X
        for i in range(1, self.L + 1):
            weight = self.weights["W{}".format(i)]
            prev_activation = self.cache["A{}".format(i-1)]
            bias = self.weights["b{}".format(i)]
            zI = np.matmul(weight, prev_activation) + bias
            self.__cache["A{}".format(i)] = 1 / (1 + np.exp(-zI))

        return self.cache["A{}".format(self.L)], self.cache

第二种初始化版本(仅__init__中的循环部分)

for i in range(1, self.L + 1):
    if i == 1:
        self.__weights["W" + str(i)] = \
            np.random.randn(layers[i - 1], nx) * np.sqrt(2 / nx)
    else:
        self.__weights["W" + str(i)] = \
            np.random.randn(layers[i - 1], layers[i - 2]) * \
            np.sqrt(2 / layers[i - 2])
    self.__weights["b" + str(i)] = np.zeros((layers[i - 1], 1))

问题表现

两种方法理论上等价:第一种先把输入特征数nx插入layers列表头部再循环;第二种则在首次循环单独处理nx逻辑。但实际测试时,第一种版本的权重/偏置形状看似正确,执行前向传播却报错。

测试代码

#!/usr/bin/env python3

import numpy as np
Deep = __import__('18-deep_neural_network').DeepNeuralNetwork

np.random.seed(18)
nx, m = np.random.randint(100, 1000, 2).tolist()
l = np.random.randint(3, 10)
sizes = np.random.randint(5, 20, l - 1).tolist()
sizes.append(1)
print(nx, sizes, m)
d = Deep(nx, sizes)
for i in range(l):
    d._DeepNeuralNetwork__weights['b' + str(i + 1)] = np.ones((sizes[i], 1))
X = np.random.randn(nx, m)
A, cache = d.forward_prop(X)
print(A)
for k, v in sorted(cache.items()):
    print(k, v)

报错信息

398 [10, 19, 1] 375
Traceback (most recent call last):
  File "18test.py", line 17, in <module>
    A, cache = d.forward_prop(X)
  File "/home/supervised_learning/classification/18-deep_neural_network.py", line 64, in forward_prop
    zI = np.matmul(weight, prev_activation) + bias
ValueError: operands could not be broadcast together with shapes (10,375) (398,1)

问题原因分析

问题出在对传入的layers列表进行了原地修改:第一种初始化中执行了layers.insert(0, nx),这会直接改变外部传入的layers列表。

看测试代码:初始化时传入的sizes是[10,19,1],执行layers.insert(0, 398)后,这个列表被修改为[398,10,19,1]。后续测试代码中循环for i in range(l)(这里l=3),执行d._DeepNeuralNetwork__weights['b' + str(i + 1)] = np.ones((sizes[i], 1))时,sizes[i]已经变成了398,10,19,而不是原来的10,19,1。

这就导致:

  • b1被设置为形状(398,1),但第一种初始化中b1原本的形状是(10,1)(对应第一层10个节点)
  • 前向传播时,zI = matmul(W1, X) + b1,matmul(W1,X)的形状是(10,375),而b1变成了(398,1),两者形状不匹配,触发广播错误。

第二种初始化版本没有修改传入的layers列表,所以测试代码中sizes保持原始值,b1被正确设置为(10,1),因此前向传播正常运行。

解决方法

第一种初始化版本中,不要原地修改传入的layers列表,而是创建一个副本:

# 替换原有的layers.insert(0, nx)
layers_copy = layers.copy()
layers_copy.insert(0, nx)
# 后续循环使用layers_copy
for i in range(1, len(layers_copy)):
    self.__weights["W{}".format(i)] = np.random.randn(
        layers_copy[i], layers_copy[i - 1]) * np.sqrt(2 / layers_copy[i - 1])
    self.__weights["b{}".format(i)] = np.zeros((layers_copy[i], 1))

这样就不会影响外部传入的列表,测试代码中的sizes保持原始值,偏置设置正确,前向传播就能正常运行。

内容的提问来源于stack exchange,提问作者Blas Hernandez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 02:23:12