You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LSTM序列标注模型训练报错:目标维度不匹配问题咨询

Hey there, let's break down this error and get your LSTM sequence labeling model up and running!

What's Causing the Error?

The core issue is a dimension mismatch between your model's output and your target data, but it's not as confusing as it sounds:

  • Your model's final TimeDistributed layer outputs a 3D tensor with shape (200, 80, n_tags) — each token has a probability distribution across all your label classes.
  • Your target y is a 2D tensor with shape (200, 80) — each token has a single label index.

While sparse_categorical_crossentropy is designed to handle this exact combination (2D targets + 3D outputs), there are a few easy fixes depending on your preference:

Fix 1: Stick with sparse_categorical_crossentropy (Memory-Efficient)

This is the recommended approach since it avoids converting your labels to one-hot encoding (saves memory, especially with large datasets). Just make sure these two details are correct:

  1. Your y array uses integer types: Keras requires label indices for sparse loss functions to be integers (not floats).
  2. n_tags is set correctly: Double-check that n_tags equals the total number of unique labels in your dataset (e.g., 3+ for BIO-style sequence labeling).

Update your code like this:

import numpy as np

# Ensure your target labels are integers
y = y.astype(np.int32)

# Optional: Use sparse_categorical_accuracy for better metric alignment
model.compile(optimizer="adam", loss="sparse_categorical_crossentropy", metrics=["sparse_categorical_accuracy"])

history = model.fit(np.array(full_embeddings), y, batch_size=32, epochs=10, verbose=1)

Fix 2: Convert y to One-Hot Encoding

If you prefer working with one-hot encoded labels (or have an older Keras version that has issues with sparse loss in sequence tasks), convert your 2D y to a 3D one-hot array:

from keras.utils import to_categorical

# Convert labels to one-hot format (shape becomes (200, 80, n_tags))
y_one_hot = to_categorical(y, num_classes=n_tags)

# Switch to categorical crossentropy loss
model.compile(optimizer="adam", loss="categorical_crossentropy", metrics=["accuracy"])

history = model.fit(np.array(full_embeddings), y_one_hot, batch_size=32, epochs=10, verbose=1)

内容的提问来源于stack exchange,提问作者Paolopast

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 07:54:07