Keras二分类复述检测模型仅预测单一标签问题排查求助
Hey there! Let's work through the two issues you're facing with your binary paraphrase detection model step by step:
1. Why Your Original Model Only Outputs a Single Label
The core problem here is using softmax activation for a single-output binary classification task:
- When you use
Dense(1, activation='softmax'), the softmax function will always output 1.0 because it normalizes outputs to sum to 1. With only one neuron, there's no other value to balance it against, so it can't distinguish between your two classes. - For binary classification with a single output neuron,
sigmoidis the correct choice—it outputs a probability between 0 and 1, which you can threshold (e.g., >0.5 = "是", <0.5 = "否") to get your final label.
2. Why All Predictions Are Identical After Switching to Sigmoid
If switching to sigmoid leads to all predictions being the same, here are the most likely fixes to try:
a. Fix the LSTM Input Shape Mismatch
Looking at your code, you've specified input_shape=(1, 120) for your LSTM layers, which is incorrect:
- The
Embeddinglayer outputs a tensor of shape(batch_size, max_length, embedding_dim)(in your case,(batch_size, 120, 120)). - The LSTM layer expects a 3D input (batch, timesteps, features), so you don't need to manually set
input_shape—Keras infers it automatically.
Update your LSTM layers to:
left_lstm = LSTM(120)(left_embedding) right_lstm = LSTM(120)(right_embedding)
b. Verify Your Label Format
Make sure your training labels are numeric values (0 or 1) instead of string labels like "是"/"否". The binary_crossentropy loss function expects numeric targets to match the sigmoid output's probability range. If you're using string labels, convert them to integers first.
c. Check Training Duration & Hyperparameters
- Your model might not have trained long enough. Try increasing the number of epochs (e.g.,
epochs=50) and adjusting the batch size (e.g.,batch_size=32) to give the model time to learn patterns in your data. - If learning is still stuck, tweak the optimizer's learning rate. For Adam, you can try a smaller rate like
optimizer=Adam(learning_rate=0.0001)if the default 0.001 is causing unstable or negligible weight updates.
d. Add Regularization to Break Symmetry
If the model's weights aren't updating, adding regularization can help break initial symmetry and prevent stagnation:
from tensorflow.keras.regularizers import l2 left_lstm = LSTM(120, kernel_regularizer=l2(0.01))(left_embedding) right_lstm = LSTM(120, kernel_regularizer=l2(0.01))(right_embedding) model_output = Dense(1, activation='sigmoid', kernel_regularizer=l2(0.01))(concat)
e. Validate Input Data Quality
Double-check that your input sequences are properly preprocessed:
- Ensure your tokenizer correctly converts phrases to integer sequences, with consistent padding/truncation to
max_length=120. - Confirm that your training data has meaningful variation between paraphrase and non-paraphrase pairs (even if class-balanced, low-quality or overly similar data can lead to uninformative predictions).
Final Corrected Code Snippet
Here's your code with the key fixes applied:
from tensorflow.keras.layers import Input, Embedding, LSTM, concatenate, Dense from tensorflow.keras.models import Model from tensorflow.keras.optimizers import Adam left_input = Input(shape=(120, )) right_input = Input(shape=(120, )) left_embedding = Embedding(vocab_size, 120, input_length=120)(left_input) right_embedding = Embedding(vocab_size, 120, input_length=120)(right_input) # Fixed LSTM layers (no manual input_shape) left_lstm = LSTM(120)(left_embedding) right_lstm = LSTM(120)(right_embedding) concat = concatenate([left_lstm, right_lstm], name='Concatenate') model_output = Dense(1, activation='sigmoid')(concat) model = Model(inputs=[left_input, right_input], outputs=model_output, name='Final_output') # Added accuracy metric to track training progress model.compile(optimizer=Adam(learning_rate=0.001), loss='binary_crossentropy', metrics=['accuracy']) model.summary()
内容的提问来源于stack exchange,提问作者Ashar

