LSTM序列标注模型训练报错:目标维度不匹配问题咨询
Hey there, let's break down this error and get your LSTM sequence labeling model up and running!
What's Causing the Error?
The core issue is a dimension mismatch between your model's output and your target data, but it's not as confusing as it sounds:
- Your model's final
TimeDistributedlayer outputs a 3D tensor with shape(200, 80, n_tags)— each token has a probability distribution across all your label classes. - Your target
yis a 2D tensor with shape(200, 80)— each token has a single label index.
While sparse_categorical_crossentropy is designed to handle this exact combination (2D targets + 3D outputs), there are a few easy fixes depending on your preference:
Fix 1: Stick with sparse_categorical_crossentropy (Memory-Efficient)
This is the recommended approach since it avoids converting your labels to one-hot encoding (saves memory, especially with large datasets). Just make sure these two details are correct:
- Your
yarray uses integer types: Keras requires label indices for sparse loss functions to be integers (not floats). n_tagsis set correctly: Double-check thatn_tagsequals the total number of unique labels in your dataset (e.g., 3+ for BIO-style sequence labeling).
Update your code like this:
import numpy as np # Ensure your target labels are integers y = y.astype(np.int32) # Optional: Use sparse_categorical_accuracy for better metric alignment model.compile(optimizer="adam", loss="sparse_categorical_crossentropy", metrics=["sparse_categorical_accuracy"]) history = model.fit(np.array(full_embeddings), y, batch_size=32, epochs=10, verbose=1)
Fix 2: Convert y to One-Hot Encoding
If you prefer working with one-hot encoded labels (or have an older Keras version that has issues with sparse loss in sequence tasks), convert your 2D y to a 3D one-hot array:
from keras.utils import to_categorical # Convert labels to one-hot format (shape becomes (200, 80, n_tags)) y_one_hot = to_categorical(y, num_classes=n_tags) # Switch to categorical crossentropy loss model.compile(optimizer="adam", loss="categorical_crossentropy", metrics=["accuracy"]) history = model.fit(np.array(full_embeddings), y_one_hot, batch_size=32, epochs=10, verbose=1)
内容的提问来源于stack exchange,提问作者Paolopast

