神经网络奇偶判断代码报ValueError:维度不匹配问题求解
问题描述
运行一段用于判断数字奇偶的神经网络代码时,持续抛出ValueError,报错信息如下:
Loss at iteration 0: 0.3568797210347673 Traceback (most recent call last): File "d:\test.py", line 102, in <module> dW1, db1, dW2, db2 = backward(X, y, a1, predictions) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "d:\Neural Network\Neural Network\test.py", line 78, in backward hidden_error = output_error.T.dot(W2) * sigmoid_derivative(a1) ^^^^^^^^^^^^^^^^^^^^^^ ValueError: shapes (1,10) and (2,1) not aligned: 10 (dim 1) != 2 (dim 0)
对应的代码:
import sqlite3 import numpy as np # Generate some training data X = np.array([[i] for i in range(10)]) y = np.array([[i % 2] for i in range(10)]) # Create a connection to the database conn = sqlite3.connect('nn.db') # Create a table to store the training data cursor = conn.cursor() cursor.execute('CREATE TABLE IF NOT EXISTS training_data (input REAL, output REAL)') # Save the training data to the database cursor.executemany('INSERT INTO training_data VALUES (?, ?)', zip(X, y)) conn.commit() # Define the neural network architecture input_size = 1 hidden_size = 2 output_size = 1 # Initialize the weights and biases randomly W1 = np.random.randn(input_size, hidden_size) b1 = np.random.randn(hidden_size) W2 = np.random.randn(hidden_size, output_size) b2 = np.random.randn(output_size) # Define the sigmoid activation function def sigmoid(x): return 1 / (1 + np.exp(-x)) # Define the derivative of the sigmoid function def sigmoid_derivative(x): return x * (1 - x) # Define the loss function def loss(predictions, targets): return np.mean((predictions - targets) ** 2) # Define the forward pass of the neural network def forward(X): # Propagate the input through the first layer z1 = X.dot(W1) + b1 a1 = sigmoid(z1) # Propagate the hidden layer output through the second layer z2 = a1.dot(W2) + b2 a2 = sigmoid(z2) return a1, a2 # Define the backward pass of the neural network def backward(X, y, a1, predictions): # Compute the error in the output layer output_error = y - predictions # Compute the gradient of the loss with respect to the output layer weights and biases dW2 = a1.T.dot(output_error * sigmoid_derivative(predictions)) db2 = np.sum(output_error * sigmoid_derivative(predictions), axis=0) # Compute the error in the hidden layer hidden_error = output_error.T.dot(W2) * sigmoid_derivative(a1) # Compute the gradient of the loss with respect to the hidden layer weights and biases dW1 = X.T.dot(hidden_error.T) db1 = np.sum(hidden_error, axis=0) return dW1, db1, dW2, db2 # Define the learning rate learning_rate = 0.1 # Train the neural network for i in range(1000): # Perform the forward pass a1, predictions = forward(X) # Compute the loss l = loss(predictions, y) # Print the loss every 100 iterations if i % 100 == 0: print(f'Loss at iteration {i}: {l}') # Perform the backward pass dW1, db1, dW2, db2 = backward(X, y, a1, predictions) # Update the weights and biases W1 += learning_rate * dW1 b1 += learning_rate * db1 W2 += learning_rate * dW2 b2 += learning_rate * db2 # Close the connection to the database conn.close() # Test the neural network on a new input test_input = np.array([[5]]) predictions = forward(test_input)[1] print(f'Prediction for test input {test_input}: {predictions}')
修复方案
报错核心是矩阵维度不匹配,反向传播中隐藏层误差的计算逻辑错误,以下是具体修改点:
1. 修正隐藏层误差计算
原代码中转置操作逻辑错误,正确计算应为输出误差乘以W2的转置,而非先转置输出误差:
hidden_error = output_error.dot(W2.T) * sigmoid_derivative(a1)
说明:output_error维度为(10,1),W2.T维度为(1,2),点乘后得到(10,2)矩阵,与a1(维度(10,2))的导数可逐元素相乘,维度完全匹配。
2. 修正隐藏层权重梯度计算
原代码中多余的转置操作需删除:
dW1 = X.T.dot(hidden_error)
说明:X.T维度为(1,10),hidden_error维度为(10,2),点乘后得到(1,2)矩阵,与W1的维度(1,2)一致,可直接用于权重更新。
3. 修正数据库存储逻辑
原代码中zip(X, y)会将数组存入数据库,导致格式错误,需提取数组内的数值:
cursor.executemany('INSERT INTO training_data VALUES (?, ?)', [(x[0], y[0]) for x, y in zip(X, y)])
完整修正后的backward函数
def backward(X, y, a1, predictions): # Compute the error in the output layer output_error = y - predictions # Compute the gradient of the loss with respect to the output layer weights and biases dW2 = a1.T.dot(output_error * sigmoid_derivative(predictions)) db2 = np.sum(output_error * sigmoid_derivative(predictions), axis=0) # Compute the error in the hidden layer hidden_error = output_error.dot(W2.T) * sigmoid_derivative(a1) # Compute the gradient of the loss with respect to the hidden layer weights and biases dW1 = X.T.dot(hidden_error) db1 = np.sum(hidden_error, axis=0) return dW1, db1, dW2, db2
内容的提问来源于stack exchange,提问作者user19533273
相关产品推荐
相关产品推荐

