搭建Sigmoid神经网络解决XOR问题时输出始终趋近0.5求助
解决XOR神经网络输出始终趋近0.5的问题
我从零开始搭建神经网络解决XOR问题,但无论输入什么数据,输出始终趋近于0.5。以下是我的实现代码:
import numpy as np import matplotlib.pyplot as plt from sklearn.model_selection import train_test_split import pandas as pd def generate_XOR_easy(): inputs = [] labels = [] for i in range(11): hasil = [round(0.1*i, 1), round(0.1*i, 1)] inputs.append(hasil) labels.append(0) if 0.1*i == 0.5: continue hasil2 = [round(0.1*i,1), round(1-0.1*i, 1)] inputs.append(hasil2) labels.append(1) return np.array(inputs), np.array(labels).reshape(21,1) def sigmoid(x): return 1.0/(1.0+np.exp(-x)) def deriv_sig(x): return np.multiply(x, 1.0-x) epochs = 1000 n_input = 2 hidden_one = 3 hidden_two = 3 n_output = 1 learning_rate = 0.001 np.random.seed(10) weigth_one = np.random.rand(n_input, hidden_one) weight_two = np.random.rand(hidden_one, hidden_two) weight_three = np.random.rand(hidden_two, n_output) xor_data = generate_XOR_easy() x_train, x_test, y_train, y_test = train_test_split(xor_data[0], xor_data[1], test_size=0.2, random_state=4) for epoch in range(epochs): #feedforward input_to_h1 = np.dot(x_train, weigth_one) output_h1 = sigmoid(input_to_h1) input_to_h2 = np.dot(output_h1, weight_two) output_h2 = sigmoid(input_to_h2) input_to_output = np.dot(output_h2, weight_three) output_output = sigmoid(input_to_output) #backpropagation error = output_output - y_train error_2_1 = error * deriv_sig(output_output) error_1_w3 = np.dot(output_h2.T, error_2_1) w3_error1 = np.dot(error_2_1, weight_three.T) h2_front = w3_error1 * deriv_sig(output_h2) h2Back_w2 = np.dot(output_h1.T, h2_front) h2Back_h1front = np.dot(h2_front, weight_two.T) h1front_h1back = h2Back_h1front * deriv_sig(output_h1) h1back_w1 = np.dot(x_train.T, h1front_h1back) weight_three -= learning_rate * error_1_w3 weight_two -= learning_rate * h2Back_w2 weigth_one -= learning_rate * h1back_w1
核心问题分析
- 数据集生成函数的
return位置错误:generate_XOR_easy的return写在for循环内部,导致循环仅执行1次就返回,最终只生成2个样本。模型在极小数据集上无法学习XOR逻辑,只能输出标签的平均概率0.5。 - 学习率过小:
0.001的学习率步长太窄,权重更新速度极慢,模型难以收敛到有效参数。 - 权重初始化不合理:
np.random.rand生成0-1的权重,易使sigmoid神经元进入饱和区,导数趋近于0,引发梯度消失,导致权重无法有效更新。 - 代码拼写与缩进错误:
weigth_one与其他weight_*变量拼写不一致;反向传播代码缩进错误,逻辑无法正常执行。
修正后的代码
import numpy as np from sklearn.model_selection import train_test_split def generate_XOR_easy(): inputs = [] labels = [] for i in range(11): hasil = [round(0.1*i, 1), round(0.1*i, 1)] inputs.append(hasil) labels.append(0) if 0.1*i == 0.5: continue hasil2 = [round(0.1*i,1), round(1-0.1*i, 1)] inputs.append(hasil2) labels.append(1) # 将return移至循环外部,生成完整21个样本 return np.array(inputs), np.array(labels).reshape(21,1) def sigmoid(x): return 1.0/(1.0+np.exp(-x)) def deriv_sig(x): return np.multiply(x, 1.0-x) epochs = 10000 # 增加迭代次数,保证收敛时间 n_input = 2 hidden_one = 3 hidden_two = 3 n_output = 1 learning_rate = 0.1 # 增大学习率,提升权重更新步长 np.random.seed(10) # 用正态分布生成小范围初始权重,避免神经元饱和 weight_one = np.random.randn(n_input, hidden_one) * 0.1 weight_two = np.random.randn(hidden_one, hidden_two) * 0.1 weight_three = np.random.randn(hidden_two, n_output) * 0.1 xor_data = generate_XOR_easy() x_train, x_test, y_train, y_test = train_test_split(xor_data[0], xor_data[1], test_size=0.2, random_state=4) for epoch in range(epochs): # feedforward input_to_h1 = np.dot(x_train, weight_one) output_h1 = sigmoid(input_to_h1) input_to_h2 = np.dot(output_h1, weight_two) output_h2 = sigmoid(input_to_h2) input_to_output = np.dot(output_h2, weight_three) output_output = sigmoid(input_to_output) # backpropagation error = output_output - y_train error_2_1 = error * deriv_sig(output_output) error_1_w3 = np.dot(output_h2.T, error_2_1) w3_error1 = np.dot(error_2_1, weight_three.T) h2_front = w3_error1 * deriv_sig(output_h2) h2Back_w2 = np.dot(output_h1.T, h2_front) h2Back_h1front = np.dot(h2_front, weight_two.T) h1front_h1back = h2Back_h1front * deriv_sig(output_h1) h1back_w1 = np.dot(x_train.T, h1front_h1back) weight_three -= learning_rate * error_1_w3 weight_two -= learning_rate * h2Back_w2 weight_one -= learning_rate * h1back_w1 # 每1000轮打印训练误差,监控收敛情况 if epoch % 1000 == 0: loss = np.mean(np.square(error)) print(f"Epoch {epoch}, Loss: {loss:.4f}") # 测试模型效果 print("\n测试集预测结果(四舍五入后):") test_output = sigmoid(np.dot(sigmoid(np.dot(sigmoid(np.dot(x_test, weight_one)), weight_two)), weight_three)) print(np.round(test_output)) print("测试集真实标签:") print(y_test)
修正说明
- 调整
generate_XOR_easy的return位置,生成完整的21个XOR样本; - 提升学习率至0.1,增加迭代次数到10000,让模型有足够步长和时间收敛;
- 使用
np.random.randn结合0.1缩放生成初始权重,避免神经元饱和; - 修正变量名拼写错误,调整反向传播代码缩进,保证逻辑正确执行。
内容的提问来源于stack exchange,提问作者ryan chandra
相关产品推荐
相关产品推荐

