You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何更新MNIST识别神经网络的第一层权重与偏置?

MNIST手写数字识别神经网络反向传播疑问

我正在构建一个不依赖深度学习库的MNIST手写数字识别神经网络,结构为:

  • 输入层:784个神经元
  • 隐藏层:10个神经元,配10个偏置
  • 输出层:10个神经元

我能理解最后一层权重的更新逻辑,但不知道如何更新第一层权重(因为最后一层的结果会影响它),也不清楚偏置的更新方法。另外如果我当前的最后一层权重更新有错误,也请指出。

之前试过模拟大量神经网络并突变最优模型,但速度太慢且没效果。


现有实现代码

#forward propagation
def forward(inp, w1, w2, biases):
    hidsRes = []
    outRes = []

    for i in range(len(w1)):
        n = np.dot(inp, w1[i])

        n += biases[i]
        n = relu(n)

        hidsRes.append(n)

    for i in range(len(w2)):
        n = np.dot(hidsRes, w2[i])

        outRes.append(n)

    return softmax(outRes)

#backpropagation
def back(avgResult, w1, w2, lr):
    for i, w in enumerate(w2):
        w2[i] += lr * avgResult[i] #I only update the last layer based on the average error of each neuron

def train(inps, hids, outs, randomWeightDiff, batchs, gens, lr):
    w1, w2, b = initNn(inps, hids, outs, randomWeightDiff)

    #loading the mnist dataset
    x_train, x_test, y_train, y_test = getData()

    for gen in range(gens):
        errors = []

        x_train, y_train = shuffle(x_train, y_train)
        
        for batch in range(batchs):  
            prediction = forward(tolist(x_train[batch].tolist()), w1, w2, b)
            y = y_train[batch]

            target = [0 if i != y else 1 for i in range(10)]

            errors.append([prediction[i] - target[i] for i in range(10)])

        print(errors)

        avg = [sum([errors[i][j] for j in range(len(errors))]) / 10 for i in range(10)]

        back(avg, w1, w2, lr)
        print("Generation {gen} \n" + f"{avg}")

train(784, 10, 10, 2, 100, 1000, 0.01)

现有代码的问题分析

1. 最后一层权重更新错误

当前的最后一层权重更新逻辑完全不成立:

  • 仅将平均误差乘以学习率直接加到权重上,忽略了隐藏层输出值(权重更新必须是误差信号与该层输入激活值的乘积)
  • 未匹配权重维度:w2作为输出层权重,维度应为[隐藏层神经元数, 输出层神经元数],现有更新方式没有对应到每个具体权重参数

2. 缺失第一层权重与偏置的更新逻辑

反向传播核心是链式法则,需要从输出层往输入层传递误差信号:

  • 通过输出层误差结合输出层权重、ReLU导数,计算隐藏层误差
  • 用隐藏层误差推导第一层权重和偏置的更新量

修正后的反向传播实现

核心步骤

  1. 计算输出层误差:output_error = prediction - target(与你当前的误差定义一致)
  2. 计算隐藏层误差:结合输出层权重、输出层误差,以及ReLU激活函数的导数
  3. 更新输出层权重:w2 += lr * np.dot(hidden_outputs.T, output_error)(批量处理时取平均)
  4. 更新隐藏层权重:w1 += lr * np.dot(inputs.T, hidden_error)
  5. 更新隐藏层偏置:biases += lr * np.mean(hidden_error, axis=0)

修正后完整代码

import numpy as np
from sklearn.utils import shuffle

def relu(x):
    return np.maximum(0, x)

def relu_derivative(x):
    return np.where(x > 0, 1, 0)

def softmax(x):
    exp_x = np.exp(x - np.max(x))  # 防止数值溢出
    return exp_x / np.sum(exp_x, axis=0)

#forward propagation
def forward(inp, w1, w2, biases):
    # inp形状:[batch_size, 784]
    hidden_pre = np.dot(inp, w1) + biases  # [batch_size, 10]
    hidden_out = relu(hidden_pre)  # [batch_size, 10]
    output_pre = np.dot(hidden_out, w2)  # [batch_size, 10]
    output_out = softmax(output_pre)  # [batch_size, 10]
    return hidden_out, output_out

#backpropagation
def back(inp, hidden_out, output_error, w1, w2, biases, lr, batch_size):
    # 计算隐藏层误差
    hidden_error = np.dot(output_error, w2.T) * relu_derivative(hidden_out)  # [batch_size, 10]
    
    # 更新输出层权重
    w2 += lr * np.dot(hidden_out.T, output_error) / batch_size  # 批量平均
    
    # 更新隐藏层权重
    w1 += lr * np.dot(inp.T, hidden_error) / batch_size
    
    # 更新隐藏层偏置
    biases += lr * np.mean(hidden_error, axis=0)

def initNn(input_size, hidden_size, output_size, weight_range):
    # 使用Xavier初始化更利于收敛
    w1 = np.random.uniform(-weight_range, weight_range, (input_size, hidden_size))
    w2 = np.random.uniform(-weight_range, weight_range, (hidden_size, output_size))
    biases = np.random.uniform(-weight_range, weight_range, hidden_size)
    return w1, w2, biases

def getData():
    # 替换为你的MNIST数据加载逻辑,这里示例用公开数据集加载方式
    from keras.datasets import mnist
    (x_train, y_train), (x_test, y_test) = mnist.load_data()
    return x_train, y_train, x_test, y_test

def train(inps, hids, outs, randomWeightDiff, batch_size, gens, lr):
    w1, w2, b = initNn(inps, hids, outs, randomWeightDiff)

    #loading the mnist dataset
    x_train, y_train, x_test, y_test = getData()
    x_train = x_train.reshape(-1, 784) / 255.0  # 归一化输入
    x_test = x_test.reshape(-1, 784) / 255.0

    total_samples = len(x_train)
    for gen in range(gens):
        x_train, y_train = shuffle(x_train, y_train)
        total_error = 0.0

        for batch_start in range(0, total_samples, batch_size):
            batch_end = batch_start + batch_size
            batch_x = x_train[batch_start:batch_end]
            batch_y = y_train[batch_start:batch_end]

            # 前向传播
            hidden_out, prediction = forward(batch_x, w1, w2, b)

            # 构建目标矩阵
            target = np.zeros((len(batch_y), outs))
            target[np.arange(len(batch_y)), batch_y] = 1

            # 计算输出层误差
            output_error = prediction - target
            total_error += np.mean(np.square(output_error))

            # 反向传播更新参数
            back(batch_x, hidden_out, output_error, w1, w2, b, lr, len(batch_x))

        # 打印每代的平均误差
        avg_error = total_error / (total_samples // batch_size)
        print(f"Generation {gen+1}, Average MSE Error: {avg_error:.4f}")

# 训练调用
train(784, 10, 10, 0.1, 100, 100, 0.01)

关键修正点说明

  • 批量处理逻辑:原代码仅处理单个样本,修正后按批量切片处理,训练效率更高
  • 权重维度匹配:确保矩阵乘法维度正确,每个权重参数都对应输入激活值与误差的乘积
  • 激活函数导数:ReLU导数是链式法则计算隐藏层误差的核心
  • 数值稳定性:Softmax计算加入最大值偏移,避免指数溢出
  • 数据归一化:输入数据归一化到0-1范围,加快模型收敛速度

内容的提问来源于stack exchange,提问作者Allo Bonjour

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 07:45:12