You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于鸢尾花数据从零实现Neural Network的维度匹配错误排查

鸢尾花分类神经网络维度不匹配问题解决

问题背景

使用numpy和pandas手动实现神经网络,基于鸢尾花数据集的4个数值特征预测3种花卉类型,运行时触发维度不匹配错误:

ValueError: operands could not be broadcast together with shapes (150,) (150,3)

报错位置在backward方法中的d_weights2 = np.dot(self.layer1.T, (2 * (y - y_hat) * d_softmax))语句。

错误原因

  1. 标签维度不匹配:原始标签y_train是字符串数组(形状(150,)),而模型输出y_hat是(150,3)的独热编码格式,两者无法直接进行算术运算。
  2. softmax导数实现错误:当前softmax_derivative返回的是每个样本的雅克比矩阵,拼接后形状为(450,450),和(150,3)的误差项相乘会导致维度混乱。
  3. 损失函数误用:当前cross_ent_loss采用的是二分类交叉熵公式,不适用于多分类任务。

解决方案

1. 标签独热编码

将字符串标签转换为独热矩阵,每个样本对应一个长度为3的one-hot向量(对应3种花卉类型)。

2. 简化梯度计算

多分类任务中,交叉熵损失对softmax输入的导数可简化为y_hat - y,无需单独计算softmax的雅克比矩阵,大幅降低复杂度并避免维度问题。

3. 替换为多分类交叉熵损失

使用针对多分类的交叉熵公式,计算每个样本的损失后取均值。

4. 修正反向传播逻辑

基于简化后的导数重新计算权重和偏置的梯度。

修改后的完整代码

import pandas as pd
import numpy as np

class NeuralNet():
    def __init__(self, i_dim, h_dim, o_dim, lr):
        self.i_dim = i_dim
        self.h_dim = h_dim
        self.o_dim = o_dim
        self.lr = lr

        self.weights1 = np.random.randn(self.i_dim, self.h_dim) / np.sqrt(self.i_dim)
        self.bias1 = np.zeros((1, self.h_dim))
        self.weights2 = np.random.randn(self.h_dim, self.o_dim) / np.sqrt(self.h_dim)
        self.bias2 = np.zeros((1, self.o_dim))

    def sigmoid(self, x):
        return 1 / (1 + np.exp(-x))

    def softmax(self, x):
        exps = np.exp(x - np.max(x, axis=1, keepdims=True))
        return exps / np.sum(exps, axis=1, keepdims=True)

    def forward(self, X):
        self.layer1 = self.sigmoid(np.dot(X, self.weights1) + self.bias1)
        self.layer2 = self.softmax(np.dot(self.layer1, self.weights2) + self.bias2)
        return self.layer2

    def sigmoid_derivative(self, x):
        return x * (1 - x)

    def backward(self, X, y, y_hat):
        # 多分类交叉熵对softmax输入的导数简化为y_hat - y
        d_loss = y_hat - y
        
        # 计算输出层梯度
        d_weights2 = np.dot(self.layer1.T, d_loss)
        d_bias2 = np.sum(d_loss, axis=0, keepdims=True)
        
        # 计算隐藏层梯度
        d_hidden = np.dot(d_loss, self.weights2.T) * self.sigmoid_derivative(self.layer1)
        d_weights1 = np.dot(X.T, d_hidden)
        d_bias1 = np.sum(d_hidden, axis=0)
        
        # 更新权重和偏置
        self.weights1 -= self.lr * d_weights1
        self.bias1 -= self.lr * d_bias1
        self.weights2 -= self.lr * d_weights2
        self.bias2 -= self.lr * d_bias2

    def cross_ent_loss(self, y, y_hat):
        # 多分类交叉熵损失,避免log(0)加小epsilon
        epsilon = 1e-10
        sample_losses = -np.sum(y * np.log(y_hat + epsilon), axis=1)
        return np.mean(sample_losses)

    def train(self, X, y, epochs):
        for epoch in range(epochs):
            y_hat = self.forward(X)
            self.backward(X, y, y_hat)
            loss = self.cross_ent_loss(y, y_hat)
            if epoch % 10 == 0:
                print(f"Epoch {epoch}: Loss = {loss:.4f}")

    def predict(self, X):
        return self.forward(X)

# 数据加载与预处理
df = pd.read_csv('/Users/brasilgu/PycharmProjects/NNfs/venv/lib/iris.data.txt', header=None)
X_train = df.iloc[:, :4].values

# 标签独热编码
y_str = df.iloc[:, -1].values
label_map = {label: idx for idx, label in enumerate(np.unique(y_str))}
y_int = np.array([label_map[label] for label in y_str])
y_train = np.eye(3)[y_int]  # 转换为(150,3)的独热矩阵

# 训练模型
nn = NeuralNet(4, 5, 3, 0.1)
nn.train(X_train, y_train, 1000)

# 预测与结果转换
y_pred = nn.predict(X_train)
y_pred_labels = np.argmax(y_pred, axis=1)
# 将预测标签转换回原始字符串
reverse_label_map = {v: k for k, v in label_map.items()}
y_pred_str = [reverse_label_map[idx] for idx in y_pred_labels]
print("预测结果(前10个):", y_pred_str[:10])

关键修改说明

  • 新增标签独热编码逻辑,将原始字符串标签转换为模型可处理的(150,3)矩阵。
  • 删除复杂的softmax_derivative方法,改用简化的梯度公式y_hat - y。
  • 替换交叉熵损失函数为多分类版本,添加epsilon避免计算log(0)的错误。
  • 重写反向传播的梯度计算逻辑,确保维度匹配。

内容的提问来源于stack exchange,提问作者philomath

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 23:15:33