You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Numpy实现的XOR问题神经网络无法收敛,求技术解决方案

Hey there, let's figure out why your XOR neural network isn't converging. Looking at the initialization code you shared, I can spot a couple of critical issues that are almost certainly holding things back:

1. 偏置初始化不合理

You initialized b1 and b2 to all-ones matrices:

b1 = np.ones((3, 1))
b2 = np.ones((1, 1))

If you're using the sigmoid activation function (the go-to choice in Andrew Ng's courses), this large initial bias will push your neurons' outputs extremely close to 1 right from the start. That lands you straight in the saturated region of the sigmoid function, where the derivative is nearly zero. When gradients are that small, gradient descent can't update your parameters effectively—so the network never learns.

Fix this by initializing biases to small random values or zeros (zeros work perfectly fine here):

b1 = np.zeros((3, 1))  # Alternatively: np.random.randn(3, 1) * 0.01
b2 = np.zeros((1, 1))

2. 权重初始化过于微小

You scaled your random weights by 0.0001:

W1 = np.random.randn(3, 2) * 0.0001
W2 = np.random.randn(1, 3) * 0.0001

While small weights are good to prevent saturation, 0.0001 is way too tiny. This will make your initial activation values barely change, resulting in extremely weak gradient signals. Training will be so slow that it looks like the network isn't converging at all.

Follow Ng's recommendation and scale by 0.01 instead:

W1 = np.random.randn(3, 2) * 0.01
W2 = np.random.randn(1, 3) * 0.01

3. 训练循环的常见疏漏(基于Ng的课程体系)

Since you didn't share your full training code, here are a few more things to check that often trip people up with XOR:

  • 激活函数选择:确保隐藏层用sigmoid(或tanh),输出层也用sigmoid(这是二分类问题)。如果用ReLU,需要调整偏置初始化避免出现“死亡神经元”。
  • 损失函数:用交叉熵损失代替均方误差(MSE)。MSE在sigmoid饱和时会出现梯度消失的问题,而交叉熵能有效缓解这个问题。
  • 学习率设置:学习率太小会导致参数更新过慢;太大则会让损失值震荡无法收敛。建议先从0.1或0.01开始尝试,再根据损失曲线调整。
  • 反向传播实现:对照Ng的符号体系仔细检查梯度计算,确保dZ1、dW1、db1等变量的推导完全正确——这里的微小错误会直接导致网络无法收敛。

修正后的初始化代码示例

Here's a fixed version of your initialization code that should set you up for success:

import numpy as np
from __future__ import print_function

X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]])
Y = np.array([[0, 1, 1, 0]])
np.random.seed(1)

# Corrected weight and bias initialization
W1 = np.random.randn(3, 2) * 0.01
b1 = np.zeros((3, 1))
W2 = np.random.randn(1, 3) * 0.01
b2 = np.zeros((1, 1))

内容的提问来源于stack exchange,提问作者Abhilash Pal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:31:01