神经网络隐藏层反向传播实现问题求助
我用C语言实现神经网络,单层网络(无隐藏层)的反向传播正常,但加入一个或多个隐藏层后就失效:首次迭代损失计算正确,但后续损失迅速飙升至无穷。
我的反向传播代码
void backpropagation(NNetwork* network) { Layer* outputLayer = network->layers[network->layerCount - 1]; // for each output for(int outputIndex = 0; outputIndex < 1; outputIndex++) { // TODO: Implement this. double totalCost = 0.0f; // the output layer's step int layerIndex = network->layerCount - 1; Layer* currentLayer = network->layers[layerIndex]; Vector* dLoss_dInputs = createVector(currentLayer->weights->rows); for(int outputNeuronIndex = 0; outputNeuronIndex < outputLayer->neuronCount; outputNeuronIndex++) { Vector* predictions = network->data->trainingOutputs->data[outputIndex]; double prediction = predictions->elements[outputNeuronIndex]; double target = network->data->yValues->elements[outputIndex]; double error = target - prediction; error *= error; error *= 0.5; /* TODO: ABSTRACT THIS TO SUIT MULTIPLE LOSS FUNCTIONS derivative of MSE is -1 * error, as: derivative of 1/2 * (value)^2 = 1/2 * 2(value) => 2/2 * (value) = 1 * value Using the chain rule for differentiation (f(g(x)) = df(g(x)) * g'(x)), we then multiply this by the derivative of the inner function, g(x) = (desired - predicted), with respect to 'predicted', which gives g'(x) = -1 Therefore, the derivative of the MSE with respect to 'predicted' is: f'(g(predicted)) * g'(predicted) = (desired - predicted) * -1 = predicted - desired */ double dLoss_dOutput = -1 * error; double dOutput_dWeightedSum = currentLayer->weightedSums->elements[outputNeuronIndex] > 0 ? 1 : 0.01; double dLoss_dWeightedSum = dLoss_dOutput * dOutput_dWeightedSum; // dLoss/dInputN = Σ [(dLoss/dOutput_i) * (dOutput_i/dWeightedSum_i) * w_i->N] // dLoss/dInputN = Σ [dLoss_dOutput * dOutput_dWeightedSum * weight] for(int weightIndex = 0; weightIndex < currentLayer->weights->rows; weightIndex++) { double dWeightedSum_dWeight = matrixToVector(network->layers[layerIndex]->input)->elements[weightIndex]; double dLoss_dWeight = dLoss_dWeightedSum * dWeightedSum_dWeight; currentLayer->gradients->data[weightIndex]->elements[outputNeuronIndex] = dLoss_dWeight; } for(int prevLayerNeuronIndex = 0; prevLayerNeuronIndex < network->layers[layerIndex - 1]->neuronCount; prevLayerNeuronIndex++) { double dLoss_dInputN = 0.0f; for(int weightColumnIndex = 0; weightColumnIndex < currentLayer->weights->columns; weightColumnIndex++) { dLoss_dInputN += (dLoss_dOutput * dOutput_dWeightedSum) * currentLayer->weights->data[prevLayerNeuronIndex]->elements[weightColumnIndex]; } printf("DLOSS_DINPUTN: %f \n", dLoss_dInputN); dLoss_dInputs->elements[prevLayerNeuronIndex] = dLoss_dInputN; } } for(layerIndex = network->layerCount - 2; layerIndex >= 0; layerIndex --) { currentLayer = network->layers[layerIndex]; Vector* dLoss_dInputsHidden = createVector(currentLayer->weights->rows); for(int neuronIndex = 0; neuronIndex < currentLayer->neuronCount; neuronIndex++) { double dLoss_dOutput = dLoss_dInputs->elements[neuronIndex]; double dOutput_dWeightedSum = currentLayer->weightedSums->elements[neuronIndex] > 0 ? 1 : 0.01; double dLoss_dWeightedSum = dLoss_dOutput * dOutput_dWeightedSum; for(int weightIndex = 0; weightIndex < currentLayer->weights->rows; weightIndex++) { double dWeightedSum_dWeight = matrixToVector(network->layers[layerIndex]->input)->elements[weightIndex]; double dLoss_dWeight = dLoss_dWeightedSum * dWeightedSum_dWeight; currentLayer->gradients->data[weightIndex]->elements[neuronIndex] = dLoss_dWeight; } if(layerIndex == 0) { continue; } for(int prevLayerNeuronIndex = 0; prevLayerNeuronIndex < network->layers[layerIndex - 1]->neuronCount; prevLayerNeuronIndex++) { double dLoss_dInputN = 0.0f; for(int weightColumnIndex = 0; weightColumnIndex < currentLayer->weights->columns; weightColumnIndex++) { dLoss_dInputN += (dLoss_dOutput * dOutput_dWeightedSum) * currentLayer->weights->data[prevLayerNeuronIndex]->elements[weightColumnIndex]; } dLoss_dInputsHidden->elements[prevLayerNeuronIndex] = dLoss_dInputN; } } freeVector(dLoss_dInputs); dLoss_dInputs = dLoss_dInputsHidden; } }}
网络结构
- 含若干神经元的输入层
- 一个或多个各含若干神经元的隐藏层
- 含单个神经元的输出层
问题排查与修复建议
1. 损失函数导数计算错误
你的注释已经推导MSE的导数应为predicted - target,但代码里错误使用了-1 * error——这里的error是已经平方并乘以0.5的MSE值,完全偏离了导数逻辑。正确的写法应为:
double error = target - prediction; double dLoss_dOutput = prediction - target; // 或等价的 -error
错误的导数会导致梯度方向完全反向,权重更新错误,直接引发损失爆炸。
2. 隐藏层误差传播的权重遍历逻辑错误
计算dLoss_dInputN时,你遍历的是currentLayer->weights->columns,但实际应该针对当前层的每个神经元,取其连接前一层神经元的权重行进行累加。正确逻辑是:对当前层第neuronIndex个神经元,它对前一层prevLayerNeuronIndex神经元的误差贡献为dLoss_dWeightedSum * currentLayer->weights->data[prevLayerNeuronIndex]->elements[neuronIndex],然后累加当前层所有神经元的该值到dLoss_dInputN。
当前的行列遍历逻辑错误,会导致误差传递混乱,梯度计算完全失效。
3. 梯度累加问题
代码中outputIndex循环仅处理单个样本,如果是批量训练,需要累加所有样本的梯度后再取平均;如果是单样本训练,当前逻辑没问题,但如果训练流程默认批量,会导致梯度仅保留最后一个样本的更新,幅度异常。
4. 学习率过高
损失爆炸的常见原因还有学习率过大,即使梯度计算正确,过大的学习率也会让权重更新幅度过大,导致损失震荡甚至飙升。可以尝试降低学习率(比如从0.1调整到0.01或0.001)测试。
5. 激活函数导数的一致性检查
你用currentLayer->weightedSums->elements[neuronIndex] > 0 ? 1 : 0.01作为Leaky ReLU的导数,逻辑本身没问题,但要确保所有隐藏层和输出层的激活函数与该导数逻辑匹配,避免部分层导数计算错误。
内容的提问来源于stack exchange,提问作者mvlcfr

