You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MNIST 0-5标签多分类模型低准确率及形状不兼容问题求助

MNIST 0-5分类模型问题排查与修复

问题原因分析

你的代码存在三个核心问题,导致准确率异常或报错:

  • 输出层神经元数量错误:当使用Dense(1)时,softmax激活后只能输出单个概率值,完全无法适配6分类任务,模型相当于在做无效的二分类尝试,因此准确率极低。
  • 标签与损失函数不匹配:使用categorical_crossentropy损失时,要求传入one-hot编码后的标签(形状为(样本数, 类别数)),但你训练时传入的是原始整数标签y_train_new(形状为(样本数,)),当把输出层改为6个神经元后,模型输出形状是(None,6),和标签形状(None,1)不兼容,触发形状错误。
  • 测试集标签处理遗漏:你直接对完整的y_test做one-hot编码,没有先筛选出0-5的样本,后续验证时会出现标签不匹配的问题。

修正后的完整代码

import tensorflow as tf
from tensorflow import keras
from keras.datasets import mnist

# 加载数据集
(x_train, y_train), (x_test, y_test) = mnist.load_data()

# 筛选0-5的训练样本
train_mask = (y_train >=0) & (y_train <=5)
x_train_new, y_train_new = x_train[train_mask], y_train[train_mask]
# 筛选0-5的测试样本
test_mask = (y_test >=0) & (y_test <=5)
x_test_new, y_test_new = x_test[test_mask], y_test[test_mask]

# 对标签做one-hot编码(类别数为6)
y_train_onehot = tf.one_hot(y_train_new, depth=6)
y_test_onehot = tf.one_hot(y_test_new, depth=6)

# 数据预处理:扁平化+归一化
x_train_final = x_train_new.reshape((-1, 784)) / 255.0
x_test_final = x_test_new.reshape((-1, 784)) / 255.0

# 构建模型:输出层6个神经元,适配6分类
model = keras.Sequential([
    keras.layers.Dense(6, activation='softmax')
])

# 编译模型:使用one-hot标签对应categorical_crossentropy损失
model.compile(
    optimizer="sgd",
    loss="categorical_crossentropy",
    metrics=["accuracy"]
)

# 训练模型:传入one-hot编码后的标签
history = model.fit(x=x_train_final, y=y_train_onehot, epochs=10, validation_data=(x_test_final, y_test_onehot))

# 打印测试集准确率
test_loss, test_acc = model.evaluate(x_test_final, y_test_onehot)
print(f"测试集准确率: {test_acc:.4f}")

额外优化说明

  • 用(y_train >=0) & (y_train <=5)替代多次或运算,代码更简洁易维护
  • 增加了验证集参数validation_data,训练时可以实时看到测试集表现
  • 训练轮次调整为10次(可根据实际情况调整),提升模型收敛效果

内容的提问来源于stack exchange,提问作者pyCaraOL

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 23:35:21