Keras函数式API构建CNN时数据维度不匹配问题咨询
Hey there! That error you're seeing is a classic case of mismatched label dimensions and loss function choice—super common when working with multi-class classification, so let's sort it out step by step.
What's Causing the Error?
Your y_train is shaped (40000, 18), which means you're using one-hot encoded labels (18 classes, each sample represented as an 18-dimensional vector with a 1 in the correct class position). But you picked sparse_categorical_crossentropy as your loss function—and this loss expects integer index labels (shape (40000,), where each value is an integer between 0 and 17 representing the class).
That's why Keras is complaining: your model's output (from dense_11) is set up for 18 classes, but the loss function is expecting a single integer per sample.
Correct Dimension Rules for Keras Functional API
Let's break down the key rules for matching your data, model, and loss function:
1. For One-Hot Encoded Labels (Your Current y_train)
- Label shape:
(num_samples, num_classes)→(40000, 18)in your case - Model output layer: Use
Dense(num_classes, activation='softmax')→ this outputs a(num_samples, 18)vector of class probabilities - Loss function: Use
categorical_crossentropy—it's designed to compare one-hot labels to softmax probabilities
2. For Integer Index Labels
If you want to use sparse_categorical_crossentropy, first convert your one-hot labels to integer indices:
import numpy as np y_train_indices = np.argmax(y_train, axis=1) # Shape becomes (40000,)
Then:
- Label shape:
(num_samples,)→(40000,) - Model output layer: Still
Dense(num_classes, activation='softmax')(output shape(num_samples, 18)) - Loss function:
sparse_categorical_crossentropy—it handles mapping integer indices to the softmax output automatically
Working Functional API Model Example (For Your Data)
Here's a complete, correct model using your one-hot labels:
from keras.layers import Input, Dense, Dropout from keras.models import Model # Define input layer: specify the shape of a single sample (no need for num_samples) input_features = Input(shape=(5371,)) # Add hidden layers (adjust sizes/dropout as needed for your task) x = Dense(256, activation='relu')(input_features) x = Dropout(0.5)(x) x = Dense(128, activation='relu')(x) # Output layer: 18 classes, softmax activation for multi-class probabilities output_classes = Dense(18, activation='softmax')(x) # Build the model model = Model(inputs=input_features, outputs=output_classes) # Compile with the correct loss function for one-hot labels model.compile( optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'] ) # Train the model model.fit(X_train, y_train, epochs=10, batch_size=32, validation_split=0.1)
Key Takeaways to Avoid This Error
- Input layer shape: Always specify the shape of a single sample (e.g.,
(5371,)for your text features)—never include the number of samples here. - Output layer alignment: The output layer's neuron count must match the number of classes (for classification).
- Loss-function label match: This is the big one—double-check that your loss function matches how you've encoded your labels:
- One-hot →
categorical_crossentropy - Integer indices →
sparse_categorical_crossentropy - Binary classification (two classes) →
binary_crossentropy(either with one-hot labels or a single output neuron + sigmoid)
- One-hot →
内容的提问来源于stack exchange,提问作者AGUY

