You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

神经网络隐藏层反向传播实现问题求助

反向传播隐藏层实现问题

我用C语言实现神经网络,单层网络(无隐藏层)的反向传播正常,但加入一个或多个隐藏层后就失效:首次迭代损失计算正确,但后续损失迅速飙升至无穷。

我的反向传播代码

void backpropagation(NNetwork* network) {
Layer* outputLayer = network->layers[network->layerCount - 1];

// for each output
for(int outputIndex = 0; outputIndex < 1; outputIndex++) {
    
    // TODO: Implement this.
    double totalCost = 0.0f;

    // the output layer's step
    int layerIndex = network->layerCount - 1;
    Layer* currentLayer = network->layers[layerIndex];
    Vector* dLoss_dInputs = createVector(currentLayer->weights->rows);

    for(int outputNeuronIndex = 0; outputNeuronIndex < outputLayer->neuronCount; outputNeuronIndex++) {
        Vector* predictions = network->data->trainingOutputs->data[outputIndex];
        double prediction = predictions->elements[outputNeuronIndex];

        double target = network->data->yValues->elements[outputIndex];
        double error =  target - prediction;
        error *= error;
        error *= 0.5;
        /* TODO: ABSTRACT THIS TO SUIT MULTIPLE LOSS FUNCTIONS  
           derivative of MSE is -1 * error, as:
           derivative of 1/2 * (value)^2 = 1/2 * 2(value) => 2/2 * (value) = 1 * value
        
           Using the chain rule for differentiation (f(g(x)) = df(g(x)) * g'(x)), we then multiply this by the derivative of the inner function,
           g(x) = (desired - predicted), with respect to 'predicted', which gives g'(x) = -1
           Therefore, the derivative of the MSE with respect to 'predicted' is: f'(g(predicted)) * g'(predicted) = (desired - predicted) * -1 = predicted - desired
        */
        double dLoss_dOutput = -1 * error;

        double dOutput_dWeightedSum = currentLayer->weightedSums->elements[outputNeuronIndex] > 0 ? 1 : 0.01;
        double dLoss_dWeightedSum = dLoss_dOutput * dOutput_dWeightedSum;

        // dLoss/dInputN = Σ [(dLoss/dOutput_i) * (dOutput_i/dWeightedSum_i) * w_i->N]
        // dLoss/dInputN = Σ [dLoss_dOutput * dOutput_dWeightedSum * weight]

        for(int weightIndex = 0; weightIndex < currentLayer->weights->rows; weightIndex++) {
            
            double dWeightedSum_dWeight = matrixToVector(network->layers[layerIndex]->input)->elements[weightIndex];
            
            double dLoss_dWeight = dLoss_dWeightedSum * dWeightedSum_dWeight;
            
            currentLayer->gradients->data[weightIndex]->elements[outputNeuronIndex] = dLoss_dWeight;                
        }

        for(int prevLayerNeuronIndex = 0; prevLayerNeuronIndex < network->layers[layerIndex - 1]->neuronCount; prevLayerNeuronIndex++) {
            double dLoss_dInputN = 0.0f;
            for(int weightColumnIndex = 0; weightColumnIndex < currentLayer->weights->columns; weightColumnIndex++) {
                dLoss_dInputN += (dLoss_dOutput * dOutput_dWeightedSum) * currentLayer->weights->data[prevLayerNeuronIndex]->elements[weightColumnIndex];
            }
            printf("DLOSS_DINPUTN: %f \n", dLoss_dInputN);
            dLoss_dInputs->elements[prevLayerNeuronIndex] = dLoss_dInputN;
        }
    }

    for(layerIndex = network->layerCount - 2; layerIndex >= 0; layerIndex --) {
        currentLayer = network->layers[layerIndex];
        Vector* dLoss_dInputsHidden = createVector(currentLayer->weights->rows);
        for(int neuronIndex = 0; neuronIndex < currentLayer->neuronCount; neuronIndex++) {
            double dLoss_dOutput = dLoss_dInputs->elements[neuronIndex];

            double dOutput_dWeightedSum = currentLayer->weightedSums->elements[neuronIndex] > 0 ? 1 : 0.01;
            double dLoss_dWeightedSum = dLoss_dOutput * dOutput_dWeightedSum;


            for(int weightIndex = 0; weightIndex < currentLayer->weights->rows; weightIndex++) {
                
                double dWeightedSum_dWeight = matrixToVector(network->layers[layerIndex]->input)->elements[weightIndex];
                
                double dLoss_dWeight = dLoss_dWeightedSum * dWeightedSum_dWeight;
                
                currentLayer->gradients->data[weightIndex]->elements[neuronIndex] = dLoss_dWeight;
                
            }

            if(layerIndex == 0) {
                continue;
            }

            for(int prevLayerNeuronIndex = 0; prevLayerNeuronIndex < network->layers[layerIndex - 1]->neuronCount; prevLayerNeuronIndex++) {
                double dLoss_dInputN = 0.0f;
                for(int weightColumnIndex = 0; weightColumnIndex < currentLayer->weights->columns; weightColumnIndex++) {
                    dLoss_dInputN += (dLoss_dOutput * dOutput_dWeightedSum) * currentLayer->weights->data[prevLayerNeuronIndex]->elements[weightColumnIndex];
                }
                
                dLoss_dInputsHidden->elements[prevLayerNeuronIndex] = dLoss_dInputN;
            }
        }
        freeVector(dLoss_dInputs);
        dLoss_dInputs = dLoss_dInputsHidden;

    }
}}

网络结构

  • 含若干神经元的输入层
  • 一个或多个各含若干神经元的隐藏层
  • 含单个神经元的输出层

问题排查与修复建议

1. 损失函数导数计算错误

你的注释已经推导MSE的导数应为predicted - target,但代码里错误使用了-1 * error——这里的error是已经平方并乘以0.5的MSE值,完全偏离了导数逻辑。正确的写法应为:

double error = target - prediction;
double dLoss_dOutput = prediction - target; // 或等价的 -error

错误的导数会导致梯度方向完全反向,权重更新错误,直接引发损失爆炸。

2. 隐藏层误差传播的权重遍历逻辑错误

计算dLoss_dInputN时,你遍历的是currentLayer->weights->columns,但实际应该针对当前层的每个神经元,取其连接前一层神经元的权重行进行累加。正确逻辑是:对当前层第neuronIndex个神经元,它对前一层prevLayerNeuronIndex神经元的误差贡献为dLoss_dWeightedSum * currentLayer->weights->data[prevLayerNeuronIndex]->elements[neuronIndex],然后累加当前层所有神经元的该值到dLoss_dInputN。

当前的行列遍历逻辑错误,会导致误差传递混乱,梯度计算完全失效。

3. 梯度累加问题

代码中outputIndex循环仅处理单个样本,如果是批量训练,需要累加所有样本的梯度后再取平均;如果是单样本训练,当前逻辑没问题,但如果训练流程默认批量,会导致梯度仅保留最后一个样本的更新,幅度异常。

4. 学习率过高

损失爆炸的常见原因还有学习率过大,即使梯度计算正确,过大的学习率也会让权重更新幅度过大,导致损失震荡甚至飙升。可以尝试降低学习率(比如从0.1调整到0.01或0.001)测试。

5. 激活函数导数的一致性检查

你用currentLayer->weightedSums->elements[neuronIndex] > 0 ? 1 : 0.01作为Leaky ReLU的导数,逻辑本身没问题,但要确保所有隐藏层和输出层的激活函数与该导数逻辑匹配,避免部分层导数计算错误。


内容的提问来源于stack exchange,提问作者mvlcfr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 02:27:01