Keras多分类DNN训练报错及第一层定义疑问
Hey there, let's get to the bottom of that error you're seeing—it's actually a common mix-up between label formats and loss functions, not an issue with your input layer (though we'll double-check that too).
The Core Problem: Loss Function vs. Label Format
Your error message says: logits shape [23,7] and labels shape [161]—here's why that's happening:
- Your
train_yis one-hot encoded (shape23×7, 7 classes for 23 samples). - But you're using
sparse_categorical_crossentropyas your loss function. This loss expects integer labels (a 1D array where each value is 0-6, shape23×1), not one-hot encoded 2D arrays. TensorFlow is flattening your 23×7 labels into a 1D array of 161 elements, which doesn't match the 23×7 logits from your output layer.
Two Simple Fixes
You can pick either solution based on whether you want to keep your one-hot labels or convert them:
Fix 1: Use the Correct Loss Function for One-Hot Labels
Swap sparse_categorical_crossentropy for categorical_crossentropy—this loss is designed for one-hot encoded labels:
model.compile(optimizer="adam", loss="categorical_crossentropy", metrics=['accuracy']) model.fit(train_x, train_y, epochs=100)
This will align your logits shape [23,7] with your labels shape [23,7], resolving the shape mismatch.
Fix 2: Convert One-Hot Labels to Integer Labels
If you prefer to stick with sparse_categorical_crossentropy, convert your one-hot train_y to a 1D array of integers using np.argmax:
import numpy as np # Convert 23×7 one-hot labels to 23×1 integer labels (0-6) train_y_int = np.argmax(train_y, axis=1) model.compile(optimizer="adam", loss="sparse_categorical_crossentropy", metrics=['accuracy']) model.fit(train_x, train_y_int, epochs=100)
Checking Your Input Layer (It's Actually Fine!)
You worried your first layer was wrong, but let's confirm:
- Your input samples have 37 features, so
input_shape=[len(words)]is correct as long aslen(words)=37(which matches your dataset dimensions). - The
unitsparameter sets the number of neurons in the first layer—usinglen(words)=37is totally valid, though you could experiment with larger values (like 64 or 128) if you want to increase model capacity later.
Quick Note on Small Dataset Size
With only 23 training samples, training for 100 epochs is likely to cause overfitting (the model will memorize your training data instead of learning general patterns). To mitigate this, try:
- Adding
Dropoutlayers to regularize the model:model = tf.keras.models.Sequential([ tf.keras.layers.Dense(units=37, input_shape=(37,)), tf.keras.layers.Dropout(0.2), # Randomly drop 20% of neurons during training tf.keras.layers.Dense(8, activation="relu"), tf.keras.layers.Dropout(0.2), tf.keras.layers.Dense(8, activation="relu"), tf.keras.layers.Dense(7, activation="softmax") ]) - Reducing the number of epochs (e.g., 20-50)
- Using L2 regularization on your dense layers
内容的提问来源于stack exchange,提问作者Amanur Rahman

