MNIST 0-5标签多分类模型低准确率及形状不兼容问题求助
MNIST 0-5分类模型问题排查与修复
问题原因分析
你的代码存在三个核心问题,导致准确率异常或报错:
- 输出层神经元数量错误:当使用
Dense(1)时,softmax激活后只能输出单个概率值,完全无法适配6分类任务,模型相当于在做无效的二分类尝试,因此准确率极低。 - 标签与损失函数不匹配:使用
categorical_crossentropy损失时,要求传入one-hot编码后的标签(形状为(样本数, 类别数)),但你训练时传入的是原始整数标签y_train_new(形状为(样本数,)),当把输出层改为6个神经元后,模型输出形状是(None,6),和标签形状(None,1)不兼容,触发形状错误。 - 测试集标签处理遗漏:你直接对完整的
y_test做one-hot编码,没有先筛选出0-5的样本,后续验证时会出现标签不匹配的问题。
修正后的完整代码
import tensorflow as tf from tensorflow import keras from keras.datasets import mnist # 加载数据集 (x_train, y_train), (x_test, y_test) = mnist.load_data() # 筛选0-5的训练样本 train_mask = (y_train >=0) & (y_train <=5) x_train_new, y_train_new = x_train[train_mask], y_train[train_mask] # 筛选0-5的测试样本 test_mask = (y_test >=0) & (y_test <=5) x_test_new, y_test_new = x_test[test_mask], y_test[test_mask] # 对标签做one-hot编码(类别数为6) y_train_onehot = tf.one_hot(y_train_new, depth=6) y_test_onehot = tf.one_hot(y_test_new, depth=6) # 数据预处理:扁平化+归一化 x_train_final = x_train_new.reshape((-1, 784)) / 255.0 x_test_final = x_test_new.reshape((-1, 784)) / 255.0 # 构建模型:输出层6个神经元,适配6分类 model = keras.Sequential([ keras.layers.Dense(6, activation='softmax') ]) # 编译模型:使用one-hot标签对应categorical_crossentropy损失 model.compile( optimizer="sgd", loss="categorical_crossentropy", metrics=["accuracy"] ) # 训练模型:传入one-hot编码后的标签 history = model.fit(x=x_train_final, y=y_train_onehot, epochs=10, validation_data=(x_test_final, y_test_onehot)) # 打印测试集准确率 test_loss, test_acc = model.evaluate(x_test_final, y_test_onehot) print(f"测试集准确率: {test_acc:.4f}")
额外优化说明
- 用
(y_train >=0) & (y_train <=5)替代多次或运算,代码更简洁易维护 - 增加了验证集参数
validation_data,训练时可以实时看到测试集表现 - 训练轮次调整为10次(可根据实际情况调整),提升模型收敛效果
内容的提问来源于stack exchange,提问作者pyCaraOL
相关产品推荐
相关产品推荐

