You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于误差反向传播的神经网络训练中损失值上升的原因排查

神经网络损失持续上升问题排查

问题背景

实现了输入层3神经元、隐藏层4神经元、输出层1神经元的神经网络,采用误差反向传播算法训练,但训练过程中损失值持续上升,最终达到0.99999444,测试输出结果不符合预期,需排查原因。

实现代码

import numpy as np
# 定义神经网络的参数
# 输入层3个,隐藏层4个,输出层1个
W2 = np.random.randn(4, 3)  # 第二层权重矩阵
B2 = np.random.randn(4, 1)  # 第二层的偏置
W3 = np.random.randn(1, 4)  # 第三层权重矩阵
B3 = np.random.randn(1, 1)


# 定义激活函数sigmoid函数
def sigmoid(X):
    return 1 / (1 + np.exp(-X))


def sigmoid_derivative(X):
    return sigmoid(X) * (1 - sigmoid(X))


# 定义神经网络的前向传播函数
def forward(X):
    # 第2层
    Z2 = np.dot(W2, X) + np.tile(B2, (1, X.shape[1]))
    A2 = sigmoid(Z2)
    # 第3层
    Z3 = np.dot(W3, A2) + np.tile(B3, (1, X.shape[1]))
    A3 = sigmoid(Z3)
    return A3


# 定义损失值
def loss(y, y_hat):
    m = len(y)
    return np.sum((y - y_hat) ** 2) / (2 * m)


# 定义损失函数的偏导数
def loss_derivative(y, y_hat):
    return y - y_hat


# 定义反向传播函数
def backward(X, y, y_hat):
    Z2 = np.dot(W2, X)  # m * 100
    A2 = sigmoid(Z2)  # m * 100
    Z3 = np.dot(W3, A2)  # k * 100

    # 第3层, J(a)*g(z3)
    delta3 = (y - y_hat) * sigmoid_derivative(Z3)  # k * 100
    # 第2层, (δ2=(W3.T)*δ3)**g(z2)
    delta2 = np.dot(W3.T, delta3) * sigmoid_derivative(Z2)  # m * 100

    # 第3层, dw3=δ3*(a2.T), db3=δ3
    dW3 = np.dot(delta3, A2.T)
    db3 = np.sum(delta3, axis=1, keepdims=True)

    # 第2层, dw2=δ2*(a1.T), db2=δ2
    dW2 = np.dot(delta2, X.T)
    db2 = np.sum(delta2, axis=1, keepdims=True)

    return dW3, dW2, db3, db2


# 定义sigmoid函数的导数
def sigmoid_grad(x):
    return sigmoid(x) * (1 - sigmoid(x))


# 训练数据
X = np.array([[0, 0, 0], [0, 0, 1], [0, 1, 1], [1, 0, 1], [1, 1, 1]])
y = np.array([[0], [0], [1], [1], [1]])

# 定义学习率
learning_rate = 0.1

# 训练神经网络
iteratorNum = 10 * 100
J_history = np.zeros(iteratorNum)
for i in range(iteratorNum):
    # 前向传播
    y_hat = forward(X.T)
    # 计算损失函数
    J_history[i] = loss(y, y_hat)
    # 反向传播
    dW3, dW2, dB3, dB2 = backward(X.T, y.T, y_hat)
    # 更新参数
    W3 -= learning_rate * dW3
    B3 -= learning_rate * dB3
    W2 -= learning_rate * dW2
    B2 -= learning_rate * dB2

print(J_history)
case1 = np.array([[0, 0, 0], [1, 1, 0]])
print(forward(case1.T))

损失历史数据(部分)

[0.99999401 0.99999403 0.99999404 0.99999406 0.99999408 0.99999409
0.99999411 0.99999413 0.99999414 0.99999416 0.99999417 0.99999419
0.99999421 0.99999422 0.99999424 0.99999425 0.99999427 0.99999428
0.9999943  0.99999432 0.99999433 0.99999435 0.99999436 0.99999438
0.99999439 0.99999441 0.99999442 0.99999444]

测试输出结果

[[0.9999975  0.99999866]]

问题原因

1. 损失函数导数符号错误

对于MSE损失函数 J = 1/(2m)Σ(y - y_hat)²,其对y_hat的偏导数应为 -(y - y_hat),但当前代码返回的是y - y_hat。这直接导致梯度方向完全反转,参数更新时不仅不降低损失,反而让损失持续增大。

2. 反向传播中Z值计算遗漏偏置

前向传播中Z2和Z3的计算包含偏置项,但backward函数中计算Z2和Z3时未加入偏置,与前向传播逻辑不一致。这会导致sigmoid_derivative(Z2)和sigmoid_derivative(Z3)计算错误,进而影响delta值的准确性,破坏反向传播的梯度计算逻辑。

3. 损失计算维度不匹配

loss函数中,y的形状为(5,1),y_hat的形状为(1,5),直接计算(y - y_hat)会触发numpy广播,导致损失值计算错误。当前损失值接近1,远高于正常MSE损失范围(正确情况下应在0~0.5之间),就是维度不匹配导致的结果。

修正方案

  1. 修正损失函数导数符号
def loss_derivative(y, y_hat):
    return -(y - y_hat)

或直接在反向传播的delta3计算中调整符号:

delta3 = -(y - y_hat) * sigmoid_derivative(Z3)
  1. 统一反向传播中Z值计算逻辑
    在backward函数中,加入偏置项:
Z2 = np.dot(W2, X) + np.tile(B2, (1, X.shape[1]))
A2 = sigmoid(Z2)
Z3 = np.dot(W3, A2) + np.tile(B3, (1, X.shape[1]))
  1. 修正损失计算的维度问题
    调整loss函数,统一输入维度:
def loss(y, y_hat):
    m = y.shape[1]
    # 确保y与y_hat维度一致
    y = y.T if y.shape[0] != y_hat.shape[0] else y
    return np.sum((y - y_hat) ** 2) / (2 * m)

内容的提问来源于stack exchange,提问作者zp hongxiuzhe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 20:44:55