Go实现MNIST神经网络:训练后精度下降的反向传播问题
Go实现MNIST神经网络训练精度先升后降+反向传播修正后精度暴跌问题
问题背景
基于Michael Nielsen《神经网络与深度学习》中的network.py,用Go语言实现了MNIST手写数字识别神经网络,当前单轮epoch精度能达到约80%,但训练过程存在异常:
- 首轮训练精度仅约20%
- 3-4轮后精度升至80%
- 后续精度逐渐下降,在批量大小50、学习率0.5的条件下,50轮后测试集精度为4683/10000
已定位问题出在反向传播算法中,核心代码如下:
func (nn *network) backprop(x *mnist.Image, y mnist.Label) ([]*mat.Dense, []*mat.Dense) { //Defines nabla_b adn nabla_w and populates them with 0s nabla_b := make([]*mat.Dense, len(nn.biases)) for i := range nn.biases { nabla_b[i] = mat.NewDense(nn.biases[i].RawMatrix().Rows, 1, nil) } nabla_w := make([]*mat.Dense, len(nn.weights)) for i := range nn.weights { nabla_w[i] = mat.NewDense(nn.weights[i].RawMatrix().Rows, nn.weights[i].RawMatrix().Cols, nil) } //Creates a usuable input matrix with each row being a pixel with a a value at column 0 pretaining to the color value from 0 - 255 a := mat.NewDense(len(x), 1, nil) for i := 0; i < len(x); i++ { val := float64(x[i]) a.Set(i, 0, val/255) } //Creates lists of *mat.Denses for storing z matrices and activation matrices zs := make([]*mat.Dense, nn.numLayers-1) activations := make([]*mat.Dense, nn.numLayers) //Stores input in activations at 0 activations[0] = a //Feedforwards stroing zs and activations for i := 0; i < len(nn.weights); i++ { weights := nn.weights[i] biases := nn.biases[i] z := mat.NewDense(weights.RawMatrix().Rows, a.RawMatrix().Cols, nil) z.Mul(weights, a) z.Add(z, biases) zs[i] = z applySigmoid := func(_, _ int, v float64) float64 { return sigmoid(v) } z.Apply(applySigmoid, z) a = z activations[i+1] = a } //Starts backpropagation by setting delta equal to the cost_derivative using the last layer of activations and label delta := nn.cost_derivative(activations[len(activations)-1], y) //Multiplies each value in the matrix by the corresponding sigmoidPrime value in the zs matrix at the last layer for i := 0; i < delta.RawMatrix().Rows; i++ { for j := 0; j < delta.RawMatrix().Cols; j++ { delta.Set(i, j, delta.At(i, j)*sigmoidPrime(zs[len(zs)-1].At(i, j))) } } //Sets the outputs last layers - bias to delta and weights to delta * last layer of activations transposed nabla_b[len(nabla_b)-1] = delta nabla_w[len(nabla_w)-1].Mul(delta, activations[len(activations)-1].T()) //For the number of layers back propagates for l := 2; l < nn.numLayers; l++ { z := zs[len(zs)-l] temp := z applySigmoidPrime := func(_, _ int, v float64) float64 { return sigmoidPrime(v) } //Applies sigmoidPrime to the z matrix stored in temp temp.Apply(applySigmoidPrime, z) //Sets delta to weights transposed * delta delta.Mul(nn.weights[len(nn.weights)-l+1].T(), delta) //Multiples the elements by the sigmoidPrime elements of the z matrix delta.MulElem(delta, temp) //Sets the outputs similarly nabla_b[len(nabla_b)-l] = delta nabla_w[len(nabla_w)-l].Mul(delta, activations[len(activations)-l-1].T()) } return nabla_b, nabla_w }
已调试发现的问题
原代码中最后一层权重更新的代码存在错误:
nabla_w[len(nabla_w)-1].Mul(delta, activations[len(activations)-1].T())
按照反向传播公式,应该使用上一层的激活值,即修改为:
nabla_w[len(nabla_w)-1].Mul(delta, activations[len(activations)-2].T())
但修改后模型精度骤降至约0.9%,进一步分析发现:
activations[len(activations)-2](倒数第二层激活值)的数值过小- 根源在于前向传播中,sigmoid函数处理较大的z值时,输出趋近于0或1,导致反向传播时权重更新幅度过小,模型无法有效学习
尝试过对数值进行强制缩放,但未解决问题。
请求帮助
希望社区能帮忙排查:
- 为什么修正反向传播的权重更新代码后,会出现倒数第二层激活值过小的问题?
- 如何调整模型或代码,解决这个问题,让模型能够稳定训练并维持较高的精度?
内容的提问来源于stack exchange,提问作者NRed12
相关产品推荐
相关产品推荐

