You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

神经网络奇偶判断代码报ValueError:维度不匹配问题求解

问题描述

运行一段用于判断数字奇偶的神经网络代码时,持续抛出ValueError,报错信息如下:

Loss at iteration 0: 0.3568797210347673
Traceback (most recent call last):
  File "d:\test.py", line 102, in <module>
    dW1, db1, dW2, db2 = backward(X, y, a1, predictions)
                         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "d:\Neural Network\Neural Network\test.py", line 78, in backward
    hidden_error = output_error.T.dot(W2) * sigmoid_derivative(a1)
                   ^^^^^^^^^^^^^^^^^^^^^^
ValueError: shapes (1,10) and (2,1) not aligned: 10 (dim 1) != 2 (dim 0)

对应的代码:

import sqlite3
import numpy as np

# Generate some training data
X = np.array([[i] for i in range(10)])
y = np.array([[i % 2] for i in range(10)])

# Create a connection to the database
conn = sqlite3.connect('nn.db')

# Create a table to store the training data
cursor = conn.cursor()
cursor.execute('CREATE TABLE IF NOT EXISTS training_data (input REAL, output REAL)')

# Save the training data to the database
cursor.executemany('INSERT INTO training_data VALUES (?, ?)', zip(X, y))
conn.commit()

# Define the neural network architecture
input_size = 1
hidden_size = 2
output_size = 1

# Initialize the weights and biases randomly
W1 = np.random.randn(input_size, hidden_size)
b1 = np.random.randn(hidden_size)
W2 = np.random.randn(hidden_size, output_size)
b2 = np.random.randn(output_size)

# Define the sigmoid activation function
def sigmoid(x):
  return 1 / (1 + np.exp(-x))

# Define the derivative of the sigmoid function
def sigmoid_derivative(x):
  return x * (1 - x)

# Define the loss function
def loss(predictions, targets):
  return np.mean((predictions - targets) ** 2)

# Define the forward pass of the neural network
def forward(X):
  # Propagate the input through the first layer
  z1 = X.dot(W1) + b1
  a1 = sigmoid(z1)

  # Propagate the hidden layer output through the second layer
  z2 = a1.dot(W2) + b2
  a2 = sigmoid(z2)

  return a1, a2

# Define the backward pass of the neural network
def backward(X, y, a1, predictions):
  # Compute the error in the output layer
  output_error = y - predictions

  # Compute the gradient of the loss with respect to the output layer weights and biases
  dW2 = a1.T.dot(output_error * sigmoid_derivative(predictions))
  db2 = np.sum(output_error * sigmoid_derivative(predictions), axis=0)

  # Compute the error in the hidden layer
  hidden_error = output_error.T.dot(W2) * sigmoid_derivative(a1)

  # Compute the gradient of the loss with respect to the hidden layer weights and biases
  dW1 = X.T.dot(hidden_error.T)
  db1 = np.sum(hidden_error, axis=0)

  return dW1, db1, dW2, db2

# Define the learning rate
learning_rate = 0.1

# Train the neural network
for i in range(1000):
  # Perform the forward pass
  a1, predictions = forward(X)

  # Compute the loss
  l = loss(predictions, y)

  # Print the loss every 100 iterations
  if i % 100 == 0:
    print(f'Loss at iteration {i}: {l}')

  # Perform the backward pass
  dW1, db1, dW2, db2 = backward(X, y, a1, predictions)

  # Update the weights and biases
  W1 += learning_rate * dW1
  b1 += learning_rate * db1
  W2 += learning_rate * dW2
  b2 += learning_rate * db2

# Close the connection to the database
conn.close()

# Test the neural network on a new input
test_input = np.array([[5]])
predictions = forward(test_input)[1]
print(f'Prediction for test input {test_input}: {predictions}')
修复方案

报错核心是矩阵维度不匹配,反向传播中隐藏层误差的计算逻辑错误,以下是具体修改点:

1. 修正隐藏层误差计算

原代码中转置操作逻辑错误,正确计算应为输出误差乘以W2的转置,而非先转置输出误差:

hidden_error = output_error.dot(W2.T) * sigmoid_derivative(a1)

说明:output_error维度为(10,1),W2.T维度为(1,2),点乘后得到(10,2)矩阵,与a1(维度(10,2))的导数可逐元素相乘,维度完全匹配。

2. 修正隐藏层权重梯度计算

原代码中多余的转置操作需删除:

dW1 = X.T.dot(hidden_error)

说明:X.T维度为(1,10),hidden_error维度为(10,2),点乘后得到(1,2)矩阵,与W1的维度(1,2)一致,可直接用于权重更新。

3. 修正数据库存储逻辑

原代码中zip(X, y)会将数组存入数据库,导致格式错误,需提取数组内的数值:

cursor.executemany('INSERT INTO training_data VALUES (?, ?)', [(x[0], y[0]) for x, y in zip(X, y)])

完整修正后的backward函数

def backward(X, y, a1, predictions):
  # Compute the error in the output layer
  output_error = y - predictions

  # Compute the gradient of the loss with respect to the output layer weights and biases
  dW2 = a1.T.dot(output_error * sigmoid_derivative(predictions))
  db2 = np.sum(output_error * sigmoid_derivative(predictions), axis=0)

  # Compute the error in the hidden layer
  hidden_error = output_error.dot(W2.T) * sigmoid_derivative(a1)

  # Compute the gradient of the loss with respect to the hidden layer weights and biases
  dW1 = X.T.dot(hidden_error)
  db1 = np.sum(hidden_error, axis=0)

  return dW1, db1, dW2, db2

内容的提问来源于stack exchange,提问作者user19533273

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 23:40:41