You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

神经网络反向传播偏置梯度计算问题求助(附MNIST代码)

问题描述

我是神经网络领域的新手,正在学习Python,尝试训练基于MNIST数据集的手写数字识别神经网络。目前遇到两个核心问题:

  1. 反向传播中计算偏置梯度时困惑:原以为偏置梯度与z(权重矩阵运算后、sigmoid激活前的输出)的梯度一致,但数值计算结果不符;
  2. 无论使用哪种偏置梯度,训练10000轮后模型准确率始终为10%,相当于随机猜测。

请帮忙排查反向传播函数中的代码问题,我的相关代码如下:

# coding: utf-8
import random
import numpy as np
from tensorflow.examples.tutorials.mnist import input_data

# Let's read the mnist dataset
mnist = input_data.read_data_sets("MNIST_data/", one_hot=True)

def Loss_function(x,y):
    #takes 2 vertical vectors, output: half of sum of squares of differences
    return 0.5*(np.dot(np.transpose(np.array(x)-np.array(y)),np.array(x)-np.array(y)))

def sigmoid(z):
    return 1.0/(1.0+np.exp((-1)*(np.array(z))))

def sigmoid_prime(z):
    # Derivative of the sigmoid
    return sigmoid(np.array(z))*(1-sigmoid(np.array(z)))

class Network(object):
    def __init__(self, sizes):
        # initialize biases and weights with random normal distr.
        # weights are indexed by target layer first
        self.num_layers = len(sizes)
        self.sizes = sizes
        self.biases = [np.random.randn(y, 1) for y in sizes[1:]]
        self.weights = [np.random.randn(y, x) for x, y in zip(sizes[:-1], sizes[1:])]

    def feedforward(self, a):
        # Return the output of the network if "a" is input
        for b, w in zip(self.biases, self.weights):
            a = sigmoid(np.dot(w, a)+b)
        return a

    def backprop(self, x, y):
        # Return tuple (nabla_b, nabla_w) representing the gradient for the cost function C_x.
        nabla_b = [np.zeros(b.shape) for b in self.biases]
        nabla_w = [np.zeros(w.shape) for w in self.weights]
        # feedforward
        activation = x
        activations = [x] # list to store all the activations, layer by layer
        zs = [] # list to store all the z vectors, layer by layer
        for b, w in zip(self.biases, self.weights):
            z = np.dot(w, activation)+b
            zs.append(z)
            activation = sigmoid(z)
            activations.append(activation)
        # backward pass
        delta = (activations[-1] - y) * sigmoid_prime(zs[-1])
        nabla_b[-1] = delta
        nabla_w[-1] = np.dot(delta, np.transpose(activations[-2]))
        # Note that the variable l in the loop below is used a little
        # differently to the notation in Chapter 2 of the book.  Here,
        # l = 1 means the last layer of neurons, l = 2 is the second-last, etc.
        # It's a renumbering of the scheme in the book, used here to take advantage
        # of the fact that Python can use negative indices in lists.
        for l in range(2, self.num_layers):
            z = zs[-l]
            sp = sigmoid_prime(z)
            delta = np.dot(np.transpose(self.weights[-l+1]), delta) * sp
            nabla_b[-l] = delta
            nabla_w[-l] = np.dot(delta, np.transpose(activations[-l-1]))
        return (nabla_b, nabla_w)

    def update_mini_batch(self, mini_batch, eta):
        # Update the network's weights and biases by applying gradient descent using backpropagation to a single mini batch.
        # The "mini_batch" is a list of tuples (x, y), and "eta" is the learning rate.
        nabla_b = [np.zeros(b.shape) for b in self.biases]
        nabla_w = [np.zeros(w.shape) for w in self.weights]
        for x, y in mini_batch:
            delta_nabla_b, delta_nabla_w = self.backprop(x, y)
            nabla_b = [nb+dnb for nb, dnb in zip(nabla_b, delta_nabla_b)]
            nabla_w = [nw+dnw for nw, dnw in zip(nabla_w, delta_nabla_w)]
        self.weights = [w-(eta/len(mini_batch))*nw for w, nw in zip(self.weights, nabla_w)]
        self.biases = [b-(eta/len(mini_batch))*nb for b, nb in zip(self.biases, nabla_b)]

    def SGD(self, training_data, epochs, mini_batch_size, eta, test_data=None):
        # Train the neural network using mini-batch stochastic gradient descent.
        if test_data: n_test = len(test_data)
        n = len(training_data)
        for j in range(epochs):
            random.shuffle(training_data)
            mini_batches = [training_data[k:k+mini_batch_size] for k in range(0, n, mini_batch_size)]
            for mini_batch in mini_batches:
                self.update_mini_batch(mini_batch, eta)
            if test_data:
                print("Epoch {0}: {1} / {2}".format(j, self.evaluate(test_data), n_test))
            else:
                print("Epoch {0} complete".format(j))

    def evaluate(self, test_data):
        # Return the number of test inputs for which the neural network outputs the correct result.
        test_results = [(np.argmax(self.feedforward(x)), np.argmax(y)) for (x, y) in test_data]
        return sum(int(x == y) for (x, y) in test_results)

if __name__ == "__main__":
    # Create a Network with 784 input neurons, 30 hidden neurons, 10 output neurons
    net = Network([784, 30, 10])
    # Train using SGD
    net.SGD(mnist.train.images, 10000, 10, 3.0, test_data=mnist.test.images)

问题排查与解答

嘿,我帮你梳理下问题,先解决你对偏置梯度的困惑,再揪出导致准确率卡在10%的核心bug,最后给你几个优化小建议。

关于偏置梯度的误解

你说的“偏置梯度和z的梯度一致”其实是对的!在反向传播里,偏置b的梯度确实等于对应层的δ(也就是你代码里的delta)。原因很简单:z = w·a + b,对b求导的话∂z/∂b=1,根据链式法则,损失函数C对b的梯度∂C/∂b = ∂C/∂z * ∂z/∂b = ∂C/∂z = δ。

那为什么你数值计算结果不符?大概率是手动计算时搞错了维度,或者数值梯度的步长选得不合适(步长太大误差大,太小会被浮点精度干扰)。等你把代码的核心bug修复后,再验证这个结论会更准确。

导致准确率10%的致命bug

先看你最后调用SGD的代码:

net.SGD(mnist.train.images, 10000, 10, 3.0, test_data=mnist.test.images)

这里犯了一个超级关键的错误:mnist.train.images是二维数组(形状(55000,784)),但你的backprop和update_mini_batch函数,都期望training_data是**(x,y)元组的列表**,其中x是(784,1)的垂直向量,y是(10,1)的one-hot向量。

你直接把整个训练集数组传进去,training_data的每个元素变成了单一行数据(不是元组),导致backprop里的矩阵运算全乱套了——相当于模型根本没在学正确的输入输出对应关系,最后只能随机输出,所以准确率卡在10%(刚好是10个数字的随机猜测概率)。

修复方法:先把训练和测试数据转换成正确的格式:

# 把训练数据转换成(x,y)元组,x和y都是垂直向量
training_data = [(x.reshape(784, 1), y.reshape(10, 1)) for x, y in zip(mnist.train.images, mnist.train.labels)]
# 测试数据同理
test_data = [(x.reshape(784, 1), y.reshape(10, 1)) for x, y in zip(mnist.test.images, mnist.test.labels)]

然后再调用SGD,而且10000轮epochs太多了,30轮左右就能看到明显效果:

net.SGD(training_data, 30, 10, 3.0, test_data=test_data)

其他优化小建议

  • 权重初始化优化:你现在用np.random.randn生成标准正态分布的权重,对于sigmoid激活函数来说,容易导致神经元饱和(z的绝对值太大,sigmoid导数接近0,梯度消失)。可以改成按输入维度的平方根缩放:
    self.weights = [np.random.randn(y, x) / np.sqrt(x) for x, y in zip(sizes[:-1], sizes[1:])]
    
    这样初始激活值会更合理,梯度传递更顺畅。
  • 损失函数替换:虽然均方误差损失能用,但分类任务用交叉熵损失效果更好,能缓解sigmoid激活下的梯度消失问题。你可以试试替换损失函数,不过先修复上面的维度问题,准确率就能飙升了。

修复后的预期效果

当你把数据格式修正后,再运行代码,应该能看到每轮epoch的准确率稳步上升:第1轮大概能到80%左右,第10轮超过90%,30轮后能接近97%,完全不会卡在10%了。

内容的提问来源于stack exchange,提问作者Filip Parker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:01:30