神经网络反向传播偏置梯度计算问题求助(附MNIST代码)
问题描述
我是神经网络领域的新手,正在学习Python,尝试训练基于MNIST数据集的手写数字识别神经网络。目前遇到两个核心问题:
- 反向传播中计算偏置梯度时困惑:原以为偏置梯度与z(权重矩阵运算后、sigmoid激活前的输出)的梯度一致,但数值计算结果不符;
- 无论使用哪种偏置梯度,训练10000轮后模型准确率始终为10%,相当于随机猜测。
请帮忙排查反向传播函数中的代码问题,我的相关代码如下:
# coding: utf-8 import random import numpy as np from tensorflow.examples.tutorials.mnist import input_data # Let's read the mnist dataset mnist = input_data.read_data_sets("MNIST_data/", one_hot=True) def Loss_function(x,y): #takes 2 vertical vectors, output: half of sum of squares of differences return 0.5*(np.dot(np.transpose(np.array(x)-np.array(y)),np.array(x)-np.array(y))) def sigmoid(z): return 1.0/(1.0+np.exp((-1)*(np.array(z)))) def sigmoid_prime(z): # Derivative of the sigmoid return sigmoid(np.array(z))*(1-sigmoid(np.array(z))) class Network(object): def __init__(self, sizes): # initialize biases and weights with random normal distr. # weights are indexed by target layer first self.num_layers = len(sizes) self.sizes = sizes self.biases = [np.random.randn(y, 1) for y in sizes[1:]] self.weights = [np.random.randn(y, x) for x, y in zip(sizes[:-1], sizes[1:])] def feedforward(self, a): # Return the output of the network if "a" is input for b, w in zip(self.biases, self.weights): a = sigmoid(np.dot(w, a)+b) return a def backprop(self, x, y): # Return tuple (nabla_b, nabla_w) representing the gradient for the cost function C_x. nabla_b = [np.zeros(b.shape) for b in self.biases] nabla_w = [np.zeros(w.shape) for w in self.weights] # feedforward activation = x activations = [x] # list to store all the activations, layer by layer zs = [] # list to store all the z vectors, layer by layer for b, w in zip(self.biases, self.weights): z = np.dot(w, activation)+b zs.append(z) activation = sigmoid(z) activations.append(activation) # backward pass delta = (activations[-1] - y) * sigmoid_prime(zs[-1]) nabla_b[-1] = delta nabla_w[-1] = np.dot(delta, np.transpose(activations[-2])) # Note that the variable l in the loop below is used a little # differently to the notation in Chapter 2 of the book. Here, # l = 1 means the last layer of neurons, l = 2 is the second-last, etc. # It's a renumbering of the scheme in the book, used here to take advantage # of the fact that Python can use negative indices in lists. for l in range(2, self.num_layers): z = zs[-l] sp = sigmoid_prime(z) delta = np.dot(np.transpose(self.weights[-l+1]), delta) * sp nabla_b[-l] = delta nabla_w[-l] = np.dot(delta, np.transpose(activations[-l-1])) return (nabla_b, nabla_w) def update_mini_batch(self, mini_batch, eta): # Update the network's weights and biases by applying gradient descent using backpropagation to a single mini batch. # The "mini_batch" is a list of tuples (x, y), and "eta" is the learning rate. nabla_b = [np.zeros(b.shape) for b in self.biases] nabla_w = [np.zeros(w.shape) for w in self.weights] for x, y in mini_batch: delta_nabla_b, delta_nabla_w = self.backprop(x, y) nabla_b = [nb+dnb for nb, dnb in zip(nabla_b, delta_nabla_b)] nabla_w = [nw+dnw for nw, dnw in zip(nabla_w, delta_nabla_w)] self.weights = [w-(eta/len(mini_batch))*nw for w, nw in zip(self.weights, nabla_w)] self.biases = [b-(eta/len(mini_batch))*nb for b, nb in zip(self.biases, nabla_b)] def SGD(self, training_data, epochs, mini_batch_size, eta, test_data=None): # Train the neural network using mini-batch stochastic gradient descent. if test_data: n_test = len(test_data) n = len(training_data) for j in range(epochs): random.shuffle(training_data) mini_batches = [training_data[k:k+mini_batch_size] for k in range(0, n, mini_batch_size)] for mini_batch in mini_batches: self.update_mini_batch(mini_batch, eta) if test_data: print("Epoch {0}: {1} / {2}".format(j, self.evaluate(test_data), n_test)) else: print("Epoch {0} complete".format(j)) def evaluate(self, test_data): # Return the number of test inputs for which the neural network outputs the correct result. test_results = [(np.argmax(self.feedforward(x)), np.argmax(y)) for (x, y) in test_data] return sum(int(x == y) for (x, y) in test_results) if __name__ == "__main__": # Create a Network with 784 input neurons, 30 hidden neurons, 10 output neurons net = Network([784, 30, 10]) # Train using SGD net.SGD(mnist.train.images, 10000, 10, 3.0, test_data=mnist.test.images)
问题排查与解答
嘿,我帮你梳理下问题,先解决你对偏置梯度的困惑,再揪出导致准确率卡在10%的核心bug,最后给你几个优化小建议。
关于偏置梯度的误解
你说的“偏置梯度和z的梯度一致”其实是对的!在反向传播里,偏置b的梯度确实等于对应层的δ(也就是你代码里的delta)。原因很简单:z = w·a + b,对b求导的话∂z/∂b=1,根据链式法则,损失函数C对b的梯度∂C/∂b = ∂C/∂z * ∂z/∂b = ∂C/∂z = δ。
那为什么你数值计算结果不符?大概率是手动计算时搞错了维度,或者数值梯度的步长选得不合适(步长太大误差大,太小会被浮点精度干扰)。等你把代码的核心bug修复后,再验证这个结论会更准确。
导致准确率10%的致命bug
先看你最后调用SGD的代码:
net.SGD(mnist.train.images, 10000, 10, 3.0, test_data=mnist.test.images)
这里犯了一个超级关键的错误:mnist.train.images是二维数组(形状(55000,784)),但你的backprop和update_mini_batch函数,都期望training_data是**(x,y)元组的列表**,其中x是(784,1)的垂直向量,y是(10,1)的one-hot向量。
你直接把整个训练集数组传进去,training_data的每个元素变成了单一行数据(不是元组),导致backprop里的矩阵运算全乱套了——相当于模型根本没在学正确的输入输出对应关系,最后只能随机输出,所以准确率卡在10%(刚好是10个数字的随机猜测概率)。
修复方法:先把训练和测试数据转换成正确的格式:
# 把训练数据转换成(x,y)元组,x和y都是垂直向量 training_data = [(x.reshape(784, 1), y.reshape(10, 1)) for x, y in zip(mnist.train.images, mnist.train.labels)] # 测试数据同理 test_data = [(x.reshape(784, 1), y.reshape(10, 1)) for x, y in zip(mnist.test.images, mnist.test.labels)]
然后再调用SGD,而且10000轮epochs太多了,30轮左右就能看到明显效果:
net.SGD(training_data, 30, 10, 3.0, test_data=test_data)
其他优化小建议
- 权重初始化优化:你现在用
np.random.randn生成标准正态分布的权重,对于sigmoid激活函数来说,容易导致神经元饱和(z的绝对值太大,sigmoid导数接近0,梯度消失)。可以改成按输入维度的平方根缩放:
这样初始激活值会更合理,梯度传递更顺畅。self.weights = [np.random.randn(y, x) / np.sqrt(x) for x, y in zip(sizes[:-1], sizes[1:])] - 损失函数替换:虽然均方误差损失能用,但分类任务用交叉熵损失效果更好,能缓解sigmoid激活下的梯度消失问题。你可以试试替换损失函数,不过先修复上面的维度问题,准确率就能飙升了。
修复后的预期效果
当你把数据格式修正后,再运行代码,应该能看到每轮epoch的准确率稳步上升:第1轮大概能到80%左右,第10轮超过90%,30轮后能接近97%,完全不会卡在10%了。
内容的提问来源于stack exchange,提问作者Filip Parker

