神经网络近似代码中的ValueError报错排查求助
ReLU神经网络逼近$\sin^2x$的形状不匹配问题修复
问题概述
编写ReLU神经网络逼近目标函数$f(x)=\sin^2x$时,运行代码出现维度不匹配的ValueError:predictions形状为(10,),而y_train形状为(100,),无法执行广播减法运算。
核心报错原因
代码对输入数据和网络参数的维度处理错误:
- 原代码将整个训练集(长度100的一维数组)作为单个输入向量传入网络,而非每个样本单独输入。
alphas参数被初始化为(10,100),与输入维度不匹配,导致后续矩阵运算输出维度错误,最终预测值仅包含10个元素,与100个样本的标签无法匹配。
修复步骤
- 调整训练数据维度:将
x_train从(100,)改为(100,1)的二维数组,让每个样本作为独立输入。 - 修正网络参数维度:
a的维度改为(1,),匹配单输入特征的场景。alphas的维度改为(10,1),每个神经元对应1个输入权重。
- 重构网络计算逻辑:调整
f_N函数中的矩阵运算,确保输出100个预测值(与样本数一致)。 - 修正梯度计算:对应调整梯度计算中的矩阵运算,保证梯度维度与参数匹配。
完整修复代码
import numpy as np import matplotlib.pyplot as plt # Define the target function f(x) = sin(x)^2 def f(x): return np.sin(x)**2 # Generate training dataset np.random.seed(0) # For reproducibility num_samples = 100 # 调整为二维数组,每个样本作为独立输入 x_train = np.linspace(0, 2*np.pi, num_samples).reshape(-1, 1) y_train = f(x_train) # Initialize coefficients input_dim = x_train.shape[1] # 输入特征数为1 a = np.random.randn(input_dim) # 匹配输入维度,shape (1,) b = np.random.randn() # 偏置项 m = 10 # 神经元数量 alphas = np.random.randn(m, input_dim) # shape (10,1),每个神经元对应1个输入权重 ts = np.random.randn(m) # ReLU阈值,shape (10,) cs = np.random.randn(m) # ReLU输出系数,shape (10,) # Implement Neural Network Approximation def relu(x): return np.maximum(0, x) def f_N(x): # x shape: (num_samples, input_dim) linear_part = np.dot(x, a) + b # shape (num_samples,) # 计算所有神经元对所有样本的激活值,shape (m, num_samples) relu_activations = relu(np.dot(alphas, x.T) - ts[:, np.newaxis]) # 加权求和得到ReLU部分输出,shape (num_samples,) relu_part = np.dot(cs, relu_activations) return linear_part + relu_part # Define Loss Function def loss_function(a, b, cs): predictions = f_N(x_train) return np.mean((predictions - y_train.flatten())**2) # Gradient Descent Optimization def gradient_descent(learning_rate=0.01, num_iterations=1000): losses = [] global a, b, cs, alphas, ts for i in range(num_iterations): # Compute gradients predictions = f_N(x_train) error = predictions - y_train.flatten() # 统一为一维数组 error = error.reshape(-1, 1) # shape (100,1) # 计算a的梯度,shape (input_dim,) grad_a = np.mean(error * x_train, axis=0) # 计算b的梯度,标量 grad_b = np.mean(error) # 计算ReLU部分的梯度 relu_input = np.dot(alphas, x_train.T) - ts[:, np.newaxis] # shape (10,100) relu_deriv = (relu_input > 0).astype(float) # ReLU导数,shape (10,100) # 计算cs的梯度,shape (m,) grad_cs = np.mean(error.T * relu_deriv, axis=1) # 更新参数 a -= learning_rate * grad_a b -= learning_rate * grad_b cs -= learning_rate * grad_cs # 裁剪cs参数,防止过大 cs = np.clip(cs, -1, 1) # 记录当前损失 current_loss = loss_function(a, b, cs) losses.append(current_loss) # 每100轮打印一次损失 if i % 100 == 0: print(f"Iteration {i}, Loss: {current_loss:.6f}") return losses # 执行梯度下降 losses = gradient_descent() # 绘制拟合结果 plt.figure(figsize=(12, 6)) plt.plot(x_train.flatten(), y_train.flatten(), label='Target function: $f(x) = \sin(x)^2$', color='blue') y_pred = f_N(x_train) plt.plot(x_train.flatten(), y_pred, label='Neural network approximation', linestyle='--', color='red') plt.title('Neural Network Approximation vs Target Function') plt.xlabel('x') plt.ylabel('y') plt.legend() plt.grid(True) plt.show() # 绘制损失曲线 plt.figure(figsize=(10, 5)) plt.plot(losses, label='Training Loss') plt.title('Training Loss over Iterations') plt.xlabel('Iteration') plt.ylabel('Loss') plt.legend() plt.grid(True) plt.show()
修复效果说明
- 调整后
predictions形状为(100,),与y_train维度匹配,可正常计算损失和梯度。 - 网络能够正确学习$\sin^2x$的曲线,训练损失会随着迭代逐步下降。
内容的提问来源于stack exchange,提问作者Siddd
相关产品推荐
相关产品推荐

