You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从零实现的神经网络Loss无下降问题求助及代码排查

问题描述

从零实现了一个神经网络,但训练过程中Loss基本没有明显下降。尝试更换Xavier权重初始化方式、调整learning_rate后仍无改善,训练Loss输出如下:

Epoch 0, Average Loss: 0.8672735163691898
Epoch 100, Average Loss: 0.6935956011113185
Epoch 200, Average Loss: 0.690694091666978
Epoch 300, Average Loss: 0.6922357305611471
Epoch 400, Average Loss: 0.6918833076884003
Epoch 500, Average Loss: 0.6909379643394351
Epoch 600, Average Loss: 0.6902891583150265
Epoch 700, Average Loss: 0.6875228090388348
Epoch 800, Average Loss: 0.6879678899764555
Epoch 900, Average Loss: 0.6670931736764081

完整实现代码:

class NeuralNetwork:
    def __init__(self, *, input_size):
        self.input_size = input_size
        self.layers = [input_size]
        self.weights = []
        self.biases = []
        self.activations = []
        
    def add_layer(self, layer_size, activation='relu'):
        self.layers.append(layer_size)
        self.activations.append(activation)
        
    def initialize_weights(self):
        self.weights = []
        self.biases = []
        for i in range(1, len(self.layers)):
            in_dim = self.layers[i-1]
            out_dim = self.layers[i]
            stddev = np.sqrt(2 / (in_dim + out_dim))
            
            weight_matrix = np.random.normal(loc=0.0, scale=stddev, size=(out_dim, in_dim))
            bias_vector = np.random.normal(loc=0.0, scale=stddev, size=(out_dim, 1))
            
            self.weights.append(weight_matrix)
            self.biases.append(bias_vector)
    
    def activate(self, Z, activation):
        if activation == 'relu':
            return np.maximum(0, Z)
        elif activation == 'tanh':
            return np.tanh(Z)
        elif activation == 'softmax':
            exp_Z = np.exp(Z - np.max(Z, axis=0, keepdims=True))
            return exp_Z / np.sum(exp_Z, axis=0, keepdims=True)
        elif activation == 'linear':
            return Z
        elif activation == 'sigmoid':
            return 1 / (1 + np.exp(-Z))
        elif activation == 'binary':
            return (Z > 0.5).astype(int)  # Binary activation for output layer
        else:
            raise ValueError(f"Unsupported activation function: {activation}")
    
    
    def activation_derivative(self, A, activation):
        if activation == 'relu':
            return (A > 0).astype(float)
        elif activation == 'tanh':
            return 1 - np.power(A, 2)
        elif activation == 'sigmoid':
            return A * (1 - A)
        elif activation == 'linear':
            return np.ones_like(A)
        elif activation == 'softmax':
            return A * (1 - A)
        elif activation == 'binary':
            return 1
        else:
            raise ValueError(f"Unsupported activation function: {activation}")
        
        
    def feed_forward(self, X):
        A = X
        activations = [A]
        for weights, bias, activation in zip(self.weights, self.biases, self.activations):
            Z = np.dot(weights, A) + bias
            A = self.activate(Z, activation)
            activations.append(A)
        return activations
    
    
    def backward_propagation(self, X, y, activations):
        dz = []
        m = X.shape[1]
        dW = []
        dB = []
        for i in reversed(range(1, len(self.layers))):
            if i == len(self.layers)-1:
                dz = activations[i] - y
            else:
                dz = np.dot(self.weights[i].T, dz) * self.activation_derivative(activations[i], self.activations[i])
            
            dw = np.dot(dz, activations[i-1].T) / m
            db = np.sum(dz, axis=1, keepdims=True) / m
            
            dW.append(dw)
            dB.append(db)
        
        return dW[::-1], dB[::-1]  # Reverse the lists to match weights/biases order

    def update_parameters(self, dW, dB, learning_rate):
        for i in range(len(self.weights)):
            self.weights[i] -= learning_rate * dW[i]
            self.biases[i] -= learning_rate * dB[i]


    def train(self, X, y, learning_rate=0.01, epochs=1000):
        m = X.shape[1]
        for epoch in range(epochs):
            total_loss = 0
            for i in range(m):
                x_sample = X[:, i:i+1]
                y_sample = y[:, i:i+1]
                
                activations = self.feed_forward(x_sample)
                dW, dB = self.backward_propagation(x_sample, y_sample, activations)
                self.update_parameters(dW, dB, learning_rate)
                
                loss = self.compute_loss(activations[-1], y_sample)
                total_loss += loss
            
            avg_loss = total_loss / m
            if epoch % 100 == 0:
                print(f'Epoch {epoch}, Average Loss: {avg_loss}')
    
    def compute_loss(self, A, y):
        m = y.shape[1]
        loss = -np.sum(y * np.log(A + 1e-8) + (1 - y) * np.log(1 - A + 1e-8)) / m
        return loss

问题排查与修复方案

1. 输出层激活函数与损失函数不匹配

你的损失函数是二元交叉熵,但如果输出层使用binary激活函数会直接导致梯度失效:

  • binary是硬阈值离散激活,输出仅为0或1,代入损失函数时虽有数值修正,但梯度传递时无法提供连续的梯度信号,导致权重更新无意义。
  • 修复:训练阶段输出层改用sigmoid激活函数,binary仅用于最终预测阶段。

2. Softmax导数实现错误

activation_derivative中softmax的导数写的是sigmoid的导数公式A*(1-A),这是错误的。softmax的正确导数结构是diag(A) - np.outer(A, A),不过如果是和交叉熵结合,输出层梯度可以简化为A - y(你当前输出层dz计算是对的),但中间层若用softmax会直接出错,需修正。

3. 反向传播激活函数索引错误

反向传播循环中self.activations[i]的索引完全错误:

  • self.activations的长度是len(self.layers)-1(对应每个隐藏层+输出层),而循环变量i是从len(self.layers)-1倒序到1,两者索引无法对应,会导致取错激活函数甚至索引越界。
  • 修复:将循环改为遍历self.activations的索引,或者把self.activations[i]改为self.activations[i-1]。修正后的反向传播代码示例:
def backward_propagation(self, X, y, activations):
    dz = None
    m = X.shape[1]
    dW = []
    dB = []
    # 遍历激活函数列表的索引,更直观
    for idx in reversed(range(len(self.activations))):
        if idx == len(self.activations)-1:
            # 输出层梯度
            dz = activations[idx+1] - y
        else:
            # 隐藏层梯度,weights[idx+1]是当前层的下一层权重
            dz = np.dot(self.weights[idx+1].T, dz) * self.activation_derivative(activations[idx+1], self.activations[idx])
        
        dw = np.dot(dz, activations[idx].T) / m
        db = np.sum(dz, axis=1, keepdims=True) / m
        
        dW.append(dw)
        dB.append(db)
    
    return dW[::-1], dB[::-1]

4. 优化策略与学习率调整

当前用的是单样本SGD,噪声大且更新效率低:

  • 可改为小批量梯度下降(每次取16/32/64个样本计算梯度)或批量梯度下降,提升训练稳定性。
  • 单样本SGD需要更大的学习率(比如0.05~0.1),原学习率0.01太小,导致权重更新幅度不足。

5. 数据预处理缺失

如果输入特征的数值范围差异较大,会导致权重更新不均衡:

  • 需对输入数据做归一化(缩放到[0,1])或标准化(均值0方差1),保证各特征对训练的贡献均衡。

内容的提问来源于stack exchange,提问作者Omar Tarek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 21:22:32