You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

搭建Sigmoid神经网络解决XOR问题时输出始终趋近0.5求助

解决XOR神经网络输出始终趋近0.5的问题

我从零开始搭建神经网络解决XOR问题,但无论输入什么数据,输出始终趋近于0.5。以下是我的实现代码:

import numpy as np
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
import pandas as pd

def generate_XOR_easy():

inputs = []
labels = []

for i in range(11):
    hasil = [round(0.1*i, 1), round(0.1*i, 1)]
    inputs.append(hasil)
    labels.append(0)

    if 0.1*i == 0.5:
        continue

    hasil2 = [round(0.1*i,1), round(1-0.1*i, 1)]
    inputs.append(hasil2)
    labels.append(1)
    return np.array(inputs), np.array(labels).reshape(21,1)

def sigmoid(x):
   return 1.0/(1.0+np.exp(-x))

def deriv_sig(x):
   return np.multiply(x, 1.0-x)

epochs = 1000
n_input = 2
hidden_one = 3
hidden_two = 3
n_output = 1
learning_rate = 0.001

np.random.seed(10)

weigth_one = np.random.rand(n_input, hidden_one)
weight_two = np.random.rand(hidden_one, hidden_two)
weight_three = np.random.rand(hidden_two, n_output)

xor_data = generate_XOR_easy()
x_train, x_test, y_train, y_test = train_test_split(xor_data[0], xor_data[1], 
test_size=0.2, random_state=4)


for epoch in range(epochs):
    
   #feedforward
   input_to_h1 = np.dot(x_train, weigth_one)
   output_h1 = sigmoid(input_to_h1)

   input_to_h2 = np.dot(output_h1, weight_two)
   output_h2 = sigmoid(input_to_h2)

   input_to_output = np.dot(output_h2, weight_three)
   output_output = sigmoid(input_to_output)

   #backpropagation

    error = output_output - y_train

    error_2_1 = error * deriv_sig(output_output)
    error_1_w3 = np.dot(output_h2.T, error_2_1)

    w3_error1 = np.dot(error_2_1, weight_three.T)
    h2_front = w3_error1 * deriv_sig(output_h2)
    h2Back_w2 = np.dot(output_h1.T, h2_front)

    h2Back_h1front = np.dot(h2_front, weight_two.T)
    h1front_h1back = h2Back_h1front * deriv_sig(output_h1)
    h1back_w1 = np.dot(x_train.T, h1front_h1back)

    weight_three -= learning_rate * error_1_w3
    weight_two -= learning_rate * h2Back_w2
    weigth_one -= learning_rate * h1back_w1

核心问题分析

  • 数据集生成函数的return位置错误:generate_XOR_easy的return写在for循环内部,导致循环仅执行1次就返回,最终只生成2个样本。模型在极小数据集上无法学习XOR逻辑,只能输出标签的平均概率0.5。
  • 学习率过小:0.001的学习率步长太窄,权重更新速度极慢,模型难以收敛到有效参数。
  • 权重初始化不合理:np.random.rand生成0-1的权重,易使sigmoid神经元进入饱和区,导数趋近于0,引发梯度消失,导致权重无法有效更新。
  • 代码拼写与缩进错误:weigth_one与其他weight_*变量拼写不一致;反向传播代码缩进错误,逻辑无法正常执行。

修正后的代码

import numpy as np
from sklearn.model_selection import train_test_split

def generate_XOR_easy():
    inputs = []
    labels = []
    for i in range(11):
        hasil = [round(0.1*i, 1), round(0.1*i, 1)]
        inputs.append(hasil)
        labels.append(0)

        if 0.1*i == 0.5:
            continue

        hasil2 = [round(0.1*i,1), round(1-0.1*i, 1)]
        inputs.append(hasil2)
        labels.append(1)
    # 将return移至循环外部,生成完整21个样本
    return np.array(inputs), np.array(labels).reshape(21,1)

def sigmoid(x):
    return 1.0/(1.0+np.exp(-x))

def deriv_sig(x):
    return np.multiply(x, 1.0-x)

epochs = 10000  # 增加迭代次数,保证收敛时间
n_input = 2
hidden_one = 3
hidden_two = 3
n_output = 1
learning_rate = 0.1  # 增大学习率,提升权重更新步长

np.random.seed(10)

# 用正态分布生成小范围初始权重,避免神经元饱和
weight_one = np.random.randn(n_input, hidden_one) * 0.1
weight_two = np.random.randn(hidden_one, hidden_two) * 0.1
weight_three = np.random.randn(hidden_two, n_output) * 0.1

xor_data = generate_XOR_easy()
x_train, x_test, y_train, y_test = train_test_split(xor_data[0], xor_data[1], 
                                                    test_size=0.2, random_state=4)

for epoch in range(epochs):
    # feedforward
    input_to_h1 = np.dot(x_train, weight_one)
    output_h1 = sigmoid(input_to_h1)

    input_to_h2 = np.dot(output_h1, weight_two)
    output_h2 = sigmoid(input_to_h2)

    input_to_output = np.dot(output_h2, weight_three)
    output_output = sigmoid(input_to_output)

    # backpropagation
    error = output_output - y_train

    error_2_1 = error * deriv_sig(output_output)
    error_1_w3 = np.dot(output_h2.T, error_2_1)

    w3_error1 = np.dot(error_2_1, weight_three.T)
    h2_front = w3_error1 * deriv_sig(output_h2)
    h2Back_w2 = np.dot(output_h1.T, h2_front)

    h2Back_h1front = np.dot(h2_front, weight_two.T)
    h1front_h1back = h2Back_h1front * deriv_sig(output_h1)
    h1back_w1 = np.dot(x_train.T, h1front_h1back)

    weight_three -= learning_rate * error_1_w3
    weight_two -= learning_rate * h2Back_w2
    weight_one -= learning_rate * h1back_w1

    # 每1000轮打印训练误差,监控收敛情况
    if epoch % 1000 == 0:
        loss = np.mean(np.square(error))
        print(f"Epoch {epoch}, Loss: {loss:.4f}")

# 测试模型效果
print("\n测试集预测结果(四舍五入后):")
test_output = sigmoid(np.dot(sigmoid(np.dot(sigmoid(np.dot(x_test, weight_one)), weight_two)), weight_three))
print(np.round(test_output))
print("测试集真实标签:")
print(y_test)

修正说明

  1. 调整generate_XOR_easy的return位置,生成完整的21个XOR样本;
  2. 提升学习率至0.1,增加迭代次数到10000,让模型有足够步长和时间收敛;
  3. 使用np.random.randn结合0.1缩放生成初始权重,避免神经元饱和;
  4. 修正变量名拼写错误,调整反向传播代码缩进,保证逻辑正确执行。

内容的提问来源于stack exchange,提问作者ryan chandra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 09:27:04