从零实现的神经网络Loss无下降问题求助及代码排查
问题描述
从零实现了一个神经网络,但训练过程中Loss基本没有明显下降。尝试更换Xavier权重初始化方式、调整learning_rate后仍无改善,训练Loss输出如下:
Epoch 0, Average Loss: 0.8672735163691898 Epoch 100, Average Loss: 0.6935956011113185 Epoch 200, Average Loss: 0.690694091666978 Epoch 300, Average Loss: 0.6922357305611471 Epoch 400, Average Loss: 0.6918833076884003 Epoch 500, Average Loss: 0.6909379643394351 Epoch 600, Average Loss: 0.6902891583150265 Epoch 700, Average Loss: 0.6875228090388348 Epoch 800, Average Loss: 0.6879678899764555 Epoch 900, Average Loss: 0.6670931736764081
完整实现代码:
class NeuralNetwork: def __init__(self, *, input_size): self.input_size = input_size self.layers = [input_size] self.weights = [] self.biases = [] self.activations = [] def add_layer(self, layer_size, activation='relu'): self.layers.append(layer_size) self.activations.append(activation) def initialize_weights(self): self.weights = [] self.biases = [] for i in range(1, len(self.layers)): in_dim = self.layers[i-1] out_dim = self.layers[i] stddev = np.sqrt(2 / (in_dim + out_dim)) weight_matrix = np.random.normal(loc=0.0, scale=stddev, size=(out_dim, in_dim)) bias_vector = np.random.normal(loc=0.0, scale=stddev, size=(out_dim, 1)) self.weights.append(weight_matrix) self.biases.append(bias_vector) def activate(self, Z, activation): if activation == 'relu': return np.maximum(0, Z) elif activation == 'tanh': return np.tanh(Z) elif activation == 'softmax': exp_Z = np.exp(Z - np.max(Z, axis=0, keepdims=True)) return exp_Z / np.sum(exp_Z, axis=0, keepdims=True) elif activation == 'linear': return Z elif activation == 'sigmoid': return 1 / (1 + np.exp(-Z)) elif activation == 'binary': return (Z > 0.5).astype(int) # Binary activation for output layer else: raise ValueError(f"Unsupported activation function: {activation}") def activation_derivative(self, A, activation): if activation == 'relu': return (A > 0).astype(float) elif activation == 'tanh': return 1 - np.power(A, 2) elif activation == 'sigmoid': return A * (1 - A) elif activation == 'linear': return np.ones_like(A) elif activation == 'softmax': return A * (1 - A) elif activation == 'binary': return 1 else: raise ValueError(f"Unsupported activation function: {activation}") def feed_forward(self, X): A = X activations = [A] for weights, bias, activation in zip(self.weights, self.biases, self.activations): Z = np.dot(weights, A) + bias A = self.activate(Z, activation) activations.append(A) return activations def backward_propagation(self, X, y, activations): dz = [] m = X.shape[1] dW = [] dB = [] for i in reversed(range(1, len(self.layers))): if i == len(self.layers)-1: dz = activations[i] - y else: dz = np.dot(self.weights[i].T, dz) * self.activation_derivative(activations[i], self.activations[i]) dw = np.dot(dz, activations[i-1].T) / m db = np.sum(dz, axis=1, keepdims=True) / m dW.append(dw) dB.append(db) return dW[::-1], dB[::-1] # Reverse the lists to match weights/biases order def update_parameters(self, dW, dB, learning_rate): for i in range(len(self.weights)): self.weights[i] -= learning_rate * dW[i] self.biases[i] -= learning_rate * dB[i] def train(self, X, y, learning_rate=0.01, epochs=1000): m = X.shape[1] for epoch in range(epochs): total_loss = 0 for i in range(m): x_sample = X[:, i:i+1] y_sample = y[:, i:i+1] activations = self.feed_forward(x_sample) dW, dB = self.backward_propagation(x_sample, y_sample, activations) self.update_parameters(dW, dB, learning_rate) loss = self.compute_loss(activations[-1], y_sample) total_loss += loss avg_loss = total_loss / m if epoch % 100 == 0: print(f'Epoch {epoch}, Average Loss: {avg_loss}') def compute_loss(self, A, y): m = y.shape[1] loss = -np.sum(y * np.log(A + 1e-8) + (1 - y) * np.log(1 - A + 1e-8)) / m return loss
问题排查与修复方案
1. 输出层激活函数与损失函数不匹配
你的损失函数是二元交叉熵,但如果输出层使用binary激活函数会直接导致梯度失效:
binary是硬阈值离散激活,输出仅为0或1,代入损失函数时虽有数值修正,但梯度传递时无法提供连续的梯度信号,导致权重更新无意义。- 修复:训练阶段输出层改用
sigmoid激活函数,binary仅用于最终预测阶段。
2. Softmax导数实现错误
activation_derivative中softmax的导数写的是sigmoid的导数公式A*(1-A),这是错误的。softmax的正确导数结构是diag(A) - np.outer(A, A),不过如果是和交叉熵结合,输出层梯度可以简化为A - y(你当前输出层dz计算是对的),但中间层若用softmax会直接出错,需修正。
3. 反向传播激活函数索引错误
反向传播循环中self.activations[i]的索引完全错误:
self.activations的长度是len(self.layers)-1(对应每个隐藏层+输出层),而循环变量i是从len(self.layers)-1倒序到1,两者索引无法对应,会导致取错激活函数甚至索引越界。- 修复:将循环改为遍历
self.activations的索引,或者把self.activations[i]改为self.activations[i-1]。修正后的反向传播代码示例:
def backward_propagation(self, X, y, activations): dz = None m = X.shape[1] dW = [] dB = [] # 遍历激活函数列表的索引,更直观 for idx in reversed(range(len(self.activations))): if idx == len(self.activations)-1: # 输出层梯度 dz = activations[idx+1] - y else: # 隐藏层梯度,weights[idx+1]是当前层的下一层权重 dz = np.dot(self.weights[idx+1].T, dz) * self.activation_derivative(activations[idx+1], self.activations[idx]) dw = np.dot(dz, activations[idx].T) / m db = np.sum(dz, axis=1, keepdims=True) / m dW.append(dw) dB.append(db) return dW[::-1], dB[::-1]
4. 优化策略与学习率调整
当前用的是单样本SGD,噪声大且更新效率低:
- 可改为小批量梯度下降(每次取16/32/64个样本计算梯度)或批量梯度下降,提升训练稳定性。
- 单样本SGD需要更大的学习率(比如0.05~0.1),原学习率0.01太小,导致权重更新幅度不足。
5. 数据预处理缺失
如果输入特征的数值范围差异较大,会导致权重更新不均衡:
- 需对输入数据做归一化(缩放到[0,1])或标准化(均值0方差1),保证各特征对训练的贡献均衡。
内容的提问来源于stack exchange,提问作者Omar Tarek
相关产品推荐
相关产品推荐

