You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras多分类DNN训练报错及第一层定义疑问

Fixing the Shape Mismatch Error in Your Multi-Class Classification NN

Hey there, let's get to the bottom of that error you're seeing—it's actually a common mix-up between label formats and loss functions, not an issue with your input layer (though we'll double-check that too).

The Core Problem: Loss Function vs. Label Format

Your error message says: logits shape [23,7] and labels shape [161]—here's why that's happening:

  • Your train_y is one-hot encoded (shape 23×7, 7 classes for 23 samples).
  • But you're using sparse_categorical_crossentropy as your loss function. This loss expects integer labels (a 1D array where each value is 0-6, shape 23×1), not one-hot encoded 2D arrays. TensorFlow is flattening your 23×7 labels into a 1D array of 161 elements, which doesn't match the 23×7 logits from your output layer.

Two Simple Fixes

You can pick either solution based on whether you want to keep your one-hot labels or convert them:

Fix 1: Use the Correct Loss Function for One-Hot Labels

Swap sparse_categorical_crossentropy for categorical_crossentropy—this loss is designed for one-hot encoded labels:

model.compile(optimizer="adam", loss="categorical_crossentropy", metrics=['accuracy'])
model.fit(train_x, train_y, epochs=100)

This will align your logits shape [23,7] with your labels shape [23,7], resolving the shape mismatch.

Fix 2: Convert One-Hot Labels to Integer Labels

If you prefer to stick with sparse_categorical_crossentropy, convert your one-hot train_y to a 1D array of integers using np.argmax:

import numpy as np
# Convert 23×7 one-hot labels to 23×1 integer labels (0-6)
train_y_int = np.argmax(train_y, axis=1)

model.compile(optimizer="adam", loss="sparse_categorical_crossentropy", metrics=['accuracy'])
model.fit(train_x, train_y_int, epochs=100)

Checking Your Input Layer (It's Actually Fine!)

You worried your first layer was wrong, but let's confirm:

  • Your input samples have 37 features, so input_shape=[len(words)] is correct as long as len(words)=37 (which matches your dataset dimensions).
  • The units parameter sets the number of neurons in the first layer—using len(words)=37 is totally valid, though you could experiment with larger values (like 64 or 128) if you want to increase model capacity later.

Quick Note on Small Dataset Size

With only 23 training samples, training for 100 epochs is likely to cause overfitting (the model will memorize your training data instead of learning general patterns). To mitigate this, try:

  • Adding Dropout layers to regularize the model:
    model = tf.keras.models.Sequential([
        tf.keras.layers.Dense(units=37, input_shape=(37,)),
        tf.keras.layers.Dropout(0.2),  # Randomly drop 20% of neurons during training
        tf.keras.layers.Dense(8, activation="relu"),
        tf.keras.layers.Dropout(0.2),
        tf.keras.layers.Dense(8, activation="relu"),
        tf.keras.layers.Dense(7, activation="softmax")
    ])
    
  • Reducing the number of epochs (e.g., 20-50)
  • Using L2 regularization on your dense layers

内容的提问来源于stack exchange,提问作者Amanur Rahman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 19:42:49