You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras简易CNN模型报错:ValueError: Shapes (None, 1) and (None, 10)不兼容

Fixing ValueError in Your Fashion MNIST CNN Training

Hey there! Let's break down the issues causing your ValueError and get your model training smoothly.

First: Missing Channel Dimension in Input Data

The Conv2D layer expects input with a channel dimension (for grayscale images, that's 1). But when you load Fashion MNIST with fashion_mnist.load_data(), the shape of your arrays is (num_samples, 28, 28)—no channel axis. This mismatch is a major cause of the ValueError.

To fix this, you need to add a channel dimension to your training, validation, and test data:

# Add channel dimension (convert to (samples, 28, 28, 1))
X_train = np.expand_dims(X_train, axis=-1)
X_val = np.expand_dims(X_val, axis=-1)
X_test = np.expand_dims(X_test, axis=-1)

Second: Mismatched Validation Labels

In your model.fit() call, you're passing y_val (integer labels) as the validation target, but your model is trained on one-hot encoded labels (y_train_ohe). This will cause a shape mismatch because the model outputs 10-dimensional vectors, but you're feeding it 1-dimensional integer labels for validation.

Update the validation data to use the one-hot encoded version:

history = model.fit(X_train, y_train_ohe, epochs=2, validation_data=(X_val, y_val_ohe))

Optional: Optimize Activation for Multi-Class Classification

While using sigmoid with categorical_crossentropy works for one-hot encoded labels, it's more common (and often more intuitive) to use softmax activation for multi-class tasks. Softmax ensures the output values sum to 1, representing a proper probability distribution across the 10 Fashion MNIST classes. If you want to stick with your current setup, that's fine—but here's how you'd adjust it if you want to follow standard practice:

keras.layers.Dense(10, activation='softmax')

Full Corrected Code

Here's your code with all the fixes applied:

import tensorflow as tf
import keras
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
import numpy as np
from keras.utils import np_utils
from keras.models import Sequential
from keras.layers import Dense, Conv2D, Flatten, MaxPooling2D

# Load data
(X_train_val, y_train_val), (X_test, y_test) = tf.keras.datasets.fashion_mnist.load_data()
X_train, X_val, y_train, y_val = train_test_split(X_train_val, y_train_val, test_size=0.15, random_state=42)

# Normalize and add channel dimension
X_train = X_train/255
X_val = X_val/255
X_test = X_test/255
X_train = np.expand_dims(X_train, axis=-1)
X_val = np.expand_dims(X_val, axis=-1)
X_test = np.expand_dims(X_test, axis=-1)

# One-hot encode labels
y_train_ohe = keras.utils.np_utils.to_categorical(y_train, 10)
y_val_ohe = keras.utils.np_utils.to_categorical(y_val, 10)
y_test_ohe = keras.utils.np_utils.to_categorical(y_test, 10)

# Build model
model = keras.models.Sequential([
    keras.layers.Conv2D(64,7, activation='relu', input_shape=(28, 28, 1)),
    keras.layers.MaxPooling2D((2)),
    keras.layers.Flatten(),
    keras.layers.Dense(100, activation='relu'),
    keras.layers.Dense(10, activation='sigmoid')  # Or switch to 'softmax'
])

# Compile and train
model.compile(loss="categorical_crossentropy", optimizer="nadam", metrics=["accuracy"])
model.summary()
history = model.fit(X_train, y_train_ohe, epochs=2, validation_data=(X_val, y_val_ohe))

These changes should resolve the ValueError and let your model train as expected. Let me know if you run into any other issues!

内容的提问来源于stack exchange,提问作者Henri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 11:04:11