Keras LSTM模型准确率极低,请求排查故障原因
Let's break down exactly why your model is hitting that 5% accuracy wall, and how to adjust it to meet both your training task and final goal of using the last hidden state as a sentence representation:
1. Your Accuracy Metric Is Completely Misaligned with the Task
First off—you’re using accuracy as a metric for a regression task, which doesn’t make sense here.
Your model is predicting continuous 300-dimensional word embeddings, not categorical word labels. Keras’ default accuracy metric is built for classification tasks (comparing predicted class IDs to true IDs). For your setup, it’s essentially checking if every single dimension of your predicted embedding matches the true embedding exactly (or within a tiny threshold)—a practically impossible bar to clear. That’s why you’re seeing ~5% accuracy: it’s a meaningless number for this task.
Fix: Drop the accuracy metric entirely and rely on mse (your loss function) to track progress. If you eventually want to predict actual words (not embeddings), reframe the task as classification (more on that below).
2. Output Layer Activation Is Breaking Your Embedding Predictions
Your code uses activation='relu' for the TimeDistributed Dense layer, but Word2Vec embeddings include both positive and negative values. ReLU clamps all negative values to 0, which means your model can never learn to predict the negative components of the true embeddings. This is a massive bottleneck for performance.
Worse, your model summary mentions a tanh layer, but your code uses ReLU—this inconsistency suggests you might have old code or a misconfigured setup. Tanh would be slightly better (it preserves negative values in the [-1,1] range), but even that compresses your embeddings, which isn’t ideal for predicting raw Word2Vec vectors.
Fix: Use a linear activation (no activation function) for your output layer:
timeOutput = TimeDistributed(keras.layers.Dense(numDimensions, activation=None), name="output")(hiddenStates)
3. Your Task Framing Might Be Off-Target
You say you’re building a "next-word prediction" model, but right now you’re predicting the next word’s embedding, not the word itself. While this is a valid task, it’s less common than standard language modeling (predicting the next word’s class from your vocabulary).
Predicting embeddings is harder to evaluate and often leads to worse downstream performance (like your sentence representation goal) because the model isn’t forced to learn sharp, discriminative features for word identity.
Fix (if you want true next-word prediction):
- Replace your output layer with a Dense layer that has size equal to your vocabulary size, using
softmaxactivation. - Use
sparse_categorical_crossentropy(if your targets are integer word IDs) orcategorical_crossentropy(if one-hot encoded) as your loss function. - Adjust your training data so the output is the integer ID of the next word (not its embedding).
4. Training Data Alignment Issues Could Be Sabotaging Learning
You mentioned your input is a sentence like "I like green apples" and output is "like green apples ". That "" needs careful handling:
- If it’s a padding token, make sure its embedding is consistent (e.g., a zero vector or dedicated padding embedding) and that you’re masking padding during training (use
mask_zero=Truein your LSTM layer to ignore padded positions). - If it’s supposed to be an "end of sentence" (EOS) token, make sure you’ve added that to your vocabulary and assigned it a valid Word2Vec embedding (or trained one from scratch if it’s not in the pre-trained set).
Without proper handling of that final position, your model might be wasting effort learning to predict a meaningless token, dragging down overall performance.
5. Your Final Goal (Sentence Representation) Needs a Task That Encourages Good Hidden States
Your end goal is to use the last LSTM hidden state as a sentence representation. Training the model to predict next-word embeddings might not be the best way to learn useful representations—standard language modeling (predicting next-word classes) is a more proven task for teaching models to capture sentence semantics.
If you stick with predicting embeddings, make sure you’re actually using the last hidden state (not the output layer’s predictions) for your sentence representation. You can extract it with a separate model:
# Create a model to extract the final hidden state as sentence representation sentence_rep_model = keras.models.Model(inputs=words, outputs=hiddenStates[:, -1, :])
内容的提问来源于stack exchange,提问作者user9242224

