使用Numpy实现Python神经网络(XOR问题)遇训练异常求助
Hey there, let's fix your XOR neural network issue! I've looked through your code and found several key problems that are causing the slow loss descent, rebound, and that frustrating "predicting the majority class" behavior. Let's break them down one by one:
1. Critical Bug in Prediction Function
Your pred function uses the global variable x instead of the current sample X[i,:] in the loop. That means every prediction is based on the last training sample you processed, not each individual input! This is why you saw all predictions looking like the majority class ratio—you were actually just predicting the same sample over and over.
2. Redundant & Error-Prone Weight Update Loops
Nested loops for weight updates are not only inefficient but also risky for index errors (like the leftover k variable in your b1 update code, which works here only because your output layer has 1 neuron). Using vectorized operations will make your code cleaner, faster, and less likely to break.
3. Too Small Learning Rate
With single-sample SGD, alpha=0.01 is way too small to make meaningful progress on the XOR problem. You need a larger learning rate (like 0.5 or 1.0) to move the weights enough to converge.
4. Unclear Loss Metric
You're printing the total cumulative loss per epoch, which gets bigger with more samples. Using average loss (total loss divided by number of samples) makes it easier to track convergence.
Here's the fixed code with all these issues addressed:
import numpy as np import matplotlib.pyplot as plt np.random.seed(565113221) def sigmoid(x): # sigmoid function return 1/(1+np.power(np.e,-x)) def sigmoid_deriv(x): # Derivative of sigmoid (used for backprop) return sigmoid(x) * (1 - sigmoid(x)) def forward(x,W1,W2,b1,b2): # feed forward a = W1.dot(x) + b1 # Combine weight multiplication and bias in one step z = sigmoid(a) b = W2.dot(z) + b2 y = sigmoid(b) return a,z,b,y def pred(X,W1,W2,b1,b2): # Fixed predict function y_pred = np.zeros((X.shape[0],1)) for i in range(X.shape[0]): x_sample = X[i,:].reshape((-1,1)) # Use current sample instead of global x _,_,_,y_pred[i] = forward(x_sample,W1,W2,b1,b2) return y_pred X = np.array([[0,0],[0,1],[1,0],[1,1]]) # features 4 * 2 Y = np.array([[0],[1],[1],[0]]) # labels 4 * 1 epsilon = 0.12 # initialize all weighs between -0.12 ~ 0.12 W1 = np.random.random((2,2)) * epsilon * 2 - epsilon # map from input to hidden b1 = np.random.random((2,1)) * epsilon * 2 - epsilon # bias1 W2 = np.random.random((1,2)) * epsilon * 2 - epsilon # map from hidden to output b2 = np.random.random((1,1)) * epsilon * 2 - epsilon # bias2 epoch = 1000 # Increased epochs to ensure convergence alpha = 0.5 # Larger learning rate for SGD for turn in range(epoch): epoch_cost = 0 for index in range(X.shape[0]): x = X[index,:].reshape((-1,1)) y = Y[index,:].reshape((-1,1)) a,z,b,y_pred = forward(x,W1,W2,b1,b2) # feed forward # Calculate cost (cross-entropy) cost = -y.dot(np.log(y_pred)) - (1-y).dot(np.log(1-y_pred)) epoch_cost += cost # Backpropagation with vectorized operations # Output layer gradients delta_out = (y_pred - y) * sigmoid_deriv(b) # Hidden layer gradients delta_hidden = W2.T.dot(delta_out) * sigmoid_deriv(a) # Update weights and biases (no loops needed!) W2 -= alpha * delta_out.dot(z.T) b2 -= alpha * delta_out W1 -= alpha * delta_hidden.dot(x.T) b1 -= alpha * delta_hidden # Print average loss every 100 epochs to avoid clutter if turn % 100 == 0: avg_cost = epoch_cost / X.shape[0] print(f'turn: {turn}, average cost: {avg_cost[0][0]:.4f}') print('\nprediction\n',pred(X,W1,W2,b1,b2)) print('ground-truth\n',Y)
Key Changes Explained:
- Fixed prediction function: Now correctly uses each sample
X[i,:]instead of the globalx. - Vectorized backprop: Replaced nested loops with matrix operations, which aligns with your neural network formulas and eliminates index errors.
- Increased learning rate:
alpha=0.5lets the model adjust weights more aggressively for SGD. - Added sigmoid derivative function: Makes backprop calculations cleaner and easier to verify.
- Print average loss: Only prints every 100 epochs to keep output readable, showing clear convergence.
- Combined bias in forward pass: Simplified the forward calculation by adding bias directly to the weighted sum.
When you run this code, you'll see the average loss drop steadily, and the final predictions will be very close to the ground truth XOR labels.
内容的提问来源于stack exchange,提问作者lyricpoem

