You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

神经网络近似代码中的ValueError报错排查求助

ReLU神经网络逼近$\sin^2x$的形状不匹配问题修复

问题概述

编写ReLU神经网络逼近目标函数$f(x)=\sin^2x$时,运行代码出现维度不匹配的ValueError:predictions形状为(10,),而y_train形状为(100,),无法执行广播减法运算。

核心报错原因

代码对输入数据和网络参数的维度处理错误:

  • 原代码将整个训练集(长度100的一维数组)作为单个输入向量传入网络,而非每个样本单独输入。
  • alphas参数被初始化为(10,100),与输入维度不匹配,导致后续矩阵运算输出维度错误,最终预测值仅包含10个元素,与100个样本的标签无法匹配。

修复步骤

  1. 调整训练数据维度:将x_train从(100,)改为(100,1)的二维数组,让每个样本作为独立输入。
  2. 修正网络参数维度:
    • a的维度改为(1,),匹配单输入特征的场景。
    • alphas的维度改为(10,1),每个神经元对应1个输入权重。
  3. 重构网络计算逻辑:调整f_N函数中的矩阵运算,确保输出100个预测值(与样本数一致)。
  4. 修正梯度计算:对应调整梯度计算中的矩阵运算,保证梯度维度与参数匹配。

完整修复代码

import numpy as np
import matplotlib.pyplot as plt

# Define the target function f(x) = sin(x)^2
def f(x):
    return np.sin(x)**2

# Generate training dataset
np.random.seed(0)  # For reproducibility
num_samples = 100
# 调整为二维数组,每个样本作为独立输入
x_train = np.linspace(0, 2*np.pi, num_samples).reshape(-1, 1)
y_train = f(x_train)

# Initialize coefficients
input_dim = x_train.shape[1]  # 输入特征数为1
a = np.random.randn(input_dim)  # 匹配输入维度,shape (1,)
b = np.random.randn()   # 偏置项
m = 10  # 神经元数量
alphas = np.random.randn(m, input_dim)  # shape (10,1),每个神经元对应1个输入权重
ts = np.random.randn(m)         # ReLU阈值,shape (10,)
cs = np.random.randn(m)         # ReLU输出系数,shape (10,)

# Implement Neural Network Approximation
def relu(x):
    return np.maximum(0, x)

def f_N(x):
    # x shape: (num_samples, input_dim)
    linear_part = np.dot(x, a) + b  # shape (num_samples,)
    # 计算所有神经元对所有样本的激活值,shape (m, num_samples)
    relu_activations = relu(np.dot(alphas, x.T) - ts[:, np.newaxis])
    # 加权求和得到ReLU部分输出,shape (num_samples,)
    relu_part = np.dot(cs, relu_activations)
    return linear_part + relu_part

# Define Loss Function
def loss_function(a, b, cs):
    predictions = f_N(x_train)
    return np.mean((predictions - y_train.flatten())**2)

# Gradient Descent Optimization
def gradient_descent(learning_rate=0.01, num_iterations=1000):
    losses = []
    global a, b, cs, alphas, ts
    for i in range(num_iterations):
        # Compute gradients
        predictions = f_N(x_train)
        error = predictions - y_train.flatten()  # 统一为一维数组
        error = error.reshape(-1, 1)  # shape (100,1)
        
        # 计算a的梯度,shape (input_dim,)
        grad_a = np.mean(error * x_train, axis=0)
        # 计算b的梯度,标量
        grad_b = np.mean(error)
        
        # 计算ReLU部分的梯度
        relu_input = np.dot(alphas, x_train.T) - ts[:, np.newaxis]  # shape (10,100)
        relu_deriv = (relu_input > 0).astype(float)  # ReLU导数,shape (10,100)
        # 计算cs的梯度,shape (m,)
        grad_cs = np.mean(error.T * relu_deriv, axis=1)
        
        # 更新参数
        a -= learning_rate * grad_a
        b -= learning_rate * grad_b
        cs -= learning_rate * grad_cs

        # 裁剪cs参数,防止过大
        cs = np.clip(cs, -1, 1)

        # 记录当前损失
        current_loss = loss_function(a, b, cs)
        losses.append(current_loss)

        # 每100轮打印一次损失
        if i % 100 == 0:
            print(f"Iteration {i}, Loss: {current_loss:.6f}")

    return losses

# 执行梯度下降
losses = gradient_descent()

# 绘制拟合结果
plt.figure(figsize=(12, 6))
plt.plot(x_train.flatten(), y_train.flatten(), label='Target function: $f(x) = \sin(x)^2$', color='blue')
y_pred = f_N(x_train)
plt.plot(x_train.flatten(), y_pred, label='Neural network approximation', linestyle='--', color='red')
plt.title('Neural Network Approximation vs Target Function')
plt.xlabel('x')
plt.ylabel('y')
plt.legend()
plt.grid(True)
plt.show()

# 绘制损失曲线
plt.figure(figsize=(10, 5))
plt.plot(losses, label='Training Loss')
plt.title('Training Loss over Iterations')
plt.xlabel('Iteration')
plt.ylabel('Loss')
plt.legend()
plt.grid(True)
plt.show()

修复效果说明

  • 调整后predictions形状为(100,),与y_train维度匹配,可正常计算损失和梯度。
  • 网络能够正确学习$\sin^2x$的曲线,训练损失会随着迭代逐步下降。

内容的提问来源于stack exchange,提问作者Siddd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 01:36:31