You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras二分类复述检测模型仅预测单一标签问题排查求助

Troubleshooting Your Paraphrase Detection Model

Hey there! Let's work through the two issues you're facing with your binary paraphrase detection model step by step:


1. Why Your Original Model Only Outputs a Single Label

The core problem here is using softmax activation for a single-output binary classification task:

  • When you use Dense(1, activation='softmax'), the softmax function will always output 1.0 because it normalizes outputs to sum to 1. With only one neuron, there's no other value to balance it against, so it can't distinguish between your two classes.
  • For binary classification with a single output neuron, sigmoid is the correct choice—it outputs a probability between 0 and 1, which you can threshold (e.g., >0.5 = "是", <0.5 = "否") to get your final label.

2. Why All Predictions Are Identical After Switching to Sigmoid

If switching to sigmoid leads to all predictions being the same, here are the most likely fixes to try:

a. Fix the LSTM Input Shape Mismatch

Looking at your code, you've specified input_shape=(1, 120) for your LSTM layers, which is incorrect:

  • The Embedding layer outputs a tensor of shape (batch_size, max_length, embedding_dim) (in your case, (batch_size, 120, 120)).
  • The LSTM layer expects a 3D input (batch, timesteps, features), so you don't need to manually set input_shape—Keras infers it automatically.

Update your LSTM layers to:

left_lstm = LSTM(120)(left_embedding)
right_lstm = LSTM(120)(right_embedding)

b. Verify Your Label Format

Make sure your training labels are numeric values (0 or 1) instead of string labels like "是"/"否". The binary_crossentropy loss function expects numeric targets to match the sigmoid output's probability range. If you're using string labels, convert them to integers first.

c. Check Training Duration & Hyperparameters

  • Your model might not have trained long enough. Try increasing the number of epochs (e.g., epochs=50) and adjusting the batch size (e.g., batch_size=32) to give the model time to learn patterns in your data.
  • If learning is still stuck, tweak the optimizer's learning rate. For Adam, you can try a smaller rate like optimizer=Adam(learning_rate=0.0001) if the default 0.001 is causing unstable or negligible weight updates.

d. Add Regularization to Break Symmetry

If the model's weights aren't updating, adding regularization can help break initial symmetry and prevent stagnation:

from tensorflow.keras.regularizers import l2

left_lstm = LSTM(120, kernel_regularizer=l2(0.01))(left_embedding)
right_lstm = LSTM(120, kernel_regularizer=l2(0.01))(right_embedding)
model_output = Dense(1, activation='sigmoid', kernel_regularizer=l2(0.01))(concat)

e. Validate Input Data Quality

Double-check that your input sequences are properly preprocessed:

  • Ensure your tokenizer correctly converts phrases to integer sequences, with consistent padding/truncation to max_length=120.
  • Confirm that your training data has meaningful variation between paraphrase and non-paraphrase pairs (even if class-balanced, low-quality or overly similar data can lead to uninformative predictions).

Final Corrected Code Snippet

Here's your code with the key fixes applied:

from tensorflow.keras.layers import Input, Embedding, LSTM, concatenate, Dense
from tensorflow.keras.models import Model
from tensorflow.keras.optimizers import Adam

left_input = Input(shape=(120, ))
right_input = Input(shape=(120, ))

left_embedding = Embedding(vocab_size, 120, input_length=120)(left_input)
right_embedding = Embedding(vocab_size, 120, input_length=120)(right_input)

# Fixed LSTM layers (no manual input_shape)
left_lstm = LSTM(120)(left_embedding)
right_lstm = LSTM(120)(right_embedding)

concat = concatenate([left_lstm, right_lstm], name='Concatenate')
model_output = Dense(1, activation='sigmoid')(concat)

model = Model(inputs=[left_input, right_input], outputs=model_output, name='Final_output')
# Added accuracy metric to track training progress
model.compile(optimizer=Adam(learning_rate=0.001), loss='binary_crossentropy', metrics=['accuracy'])
model.summary()

内容的提问来源于stack exchange,提问作者Ashar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 18:17:39