You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多分类任务中model.fit执行报错问题排查求助

MNIST多分类CNN模型训练报错解决

问题原因

报错核心提示:logits和labels的第一维度必须相同,当前logits形状为[128,10],labels形状为[1280]
你在代码中对标签做了to_categorical处理(将整数标签转为one-hot编码的二维数组,形状为[样本数, 10]),但模型编译时使用的损失函数是sparse_categorical_crossentropy——这个损失函数要求标签是整数型的单维度数组(形状为[样本数]),两者格式不匹配导致维度错误。

解决方案(二选一即可)

方案1:更换损失函数

既然标签已经做了one-hot编码,将损失函数改为categorical_crossentropy:

model.compile(loss='categorical_crossentropy',
              optimizer='adam',
              metrics=['accuracy'])

方案2:取消标签的one-hot编码

保留原始整数标签,不使用to_categorical处理,继续使用sparse_categorical_crossentropy:
删除以下两行代码:

y_train = to_categorical(y_train, num_classes)
y_test = to_categorical(y_test, num_classes)

修正后完整代码(以方案1为例)

from keras.datasets import mnist
import matplotlib.pyplot as plt

(x_train, y_train), (x_test, y_test) = mnist.load_data()

# save input image dimensions
img_rows, img_cols = 28, 28

x_train = x_train.reshape(x_train.shape[0], img_rows, img_cols, 1)
x_test = x_test.reshape(x_test.shape[0], img_rows, img_cols, 1)

x_train = x_train / 255.0
x_test = x_test / 255.0

from keras.utils import to_categorical
num_classes = 10

y_train = to_categorical(y_train, num_classes)
y_test = to_categorical(y_test, num_classes)

from keras.models import Sequential
from keras.layers import Dense, Dropout, Flatten, Conv2D, MaxPooling2D

model = Sequential()
model.add(Conv2D(32, kernel_size=(3, 3),
     activation='relu',
     input_shape=(img_rows, img_cols, 1)))

model.add(Conv2D(64, (3, 3), activation='relu'))
model.add(MaxPooling2D(pool_size=(2, 2)))
model.add(Dropout(0.25))
model.add(Flatten())
model.add(Dense(128, activation='relu'))
model.add(Dropout(0.5))
model.add(Dense(num_classes, activation='softmax'))

# 修改损失函数为categorical_crossentropy
model.compile(loss='categorical_crossentropy',
      optimizer='adam',
      metrics=['accuracy'])

batch_size = 128
epochs = 10

model.fit(x_train, y_train,
          batch_size=batch_size,
          epochs=epochs,
          verbose=1,
          validation_data=(x_test, y_test))

score = model.evaluate(x_test, y_test, verbose=0)
print('Test loss:', score[0])
print('Test accuracy:', score[1])
model.save("test_model.h5")

import imageio
import numpy as np
from matplotlib import pyplot as plt

im = imageio.imread("https://i.imgur.com/a3Rql9C.png")

gray = np.dot(im[...,:3], [0.299, 0.587, 0.114])
plt.imshow(gray, cmap = plt.get_cmap('gray'))
plt.show()
 
# reshape the image
gray = gray.reshape(1, img_rows, img_cols, 1)

# normalize image
gray /= 255

# load the model
from keras.models import load_model
model = load_model("test_model.h5")

# predict digit
prediction = model.predict(gray)
print(prediction.argmax())

内容的提问来源于stack exchange,提问作者keith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 19:45:00