基于Keras计算两个扁平化输出的余弦距离及模型实现咨询
Alright, let's swap out that sum+softmax comparison with cosine distance for your question-answer LSTM model. Here's a straightforward way to build this into your Keras setup step by step:
1. Get your flattened encoder outputs
You already have your question encoder (qenc) set up. Let's assume you have a mirrored answer encoder (aenc). After pulling the final outputs from both LSTMs, we'll flatten them into dense vectors ready for distance calculation:
from keras.models import Sequential, Model from keras.layers import Embedding, Bidirectional, LSTM, Flatten, Lambda from keras import backend as K # Your existing parameters (tweak to match your actual setup) WORD2VEC_EMBED_SIZE = 128 vocab_size = 10000 seq_maxlen = 50 QA_EMBED_SIZE = 64 embedding_weights = ... # Your pre-trained word vector weights # Question encoder (extended from your existing code) qenc = Sequential() qenc.add(Embedding(output_dim=WORD2VEC_EMBED_SIZE, input_dim=vocab_size, input_length=seq_maxlen, weights=[embedding_weights])) qenc.add(Bidirectional(LSTM(QA_EMBED_SIZE, return_sequences=True))) # Flatten the LSTM's sequence output to a single dense vector q_flat = Flatten()(qenc.output) # Answer encoder (mirror the question encoder's structure) aenc = Sequential() aenc.add(Embedding(output_dim=WORD2VEC_EMBED_SIZE, input_dim=vocab_size, input_length=seq_maxlen, weights=[embedding_weights])) aenc.add(Bidirectional(LSTM(QA_EMBED_SIZE, return_sequences=True))) a_flat = Flatten()(aenc.output)
2. Define the cosine distance calculation
Cosine distance is 1 - cosine similarity, where similarity is the dot product of two L2-normalized vectors. We'll wrap this logic in a Lambda layer (Keras requires operations to be layer-based):
def cosine_distance_fn(inputs): q_vec, a_vec = inputs # Normalize both vectors to unit length first q_norm = K.l2_normalize(q_vec, axis=-1) a_norm = K.l2_normalize(a_vec, axis=-1) # Calculate 1 minus the dot product (dot product of normalized vectors = cosine similarity) return 1 - K.sum(q_norm * a_norm, axis=-1, keepdims=True) # Apply the function to our flattened question/answer vectors cosine_dist = Lambda(cosine_distance_fn)([q_flat, a_flat])
3. Build and compile the full model
Now we'll create a combined model that takes both question and answer inputs, and outputs the cosine distance between their encoded vectors:
# Assemble the full model model = Model(inputs=[qenc.input, aenc.input], outputs=cosine_dist) # Compile for your task (adjust optimizer/loss based on your goals) # For example, use MSE if training on matching question-answer pairs model.compile(optimizer='adam', loss='mse')
Key Tips:
- Dimension Alignment: Since both encoders have identical structures, their flattened vectors will automatically have the same dimension—critical for valid dot product calculation.
- Normalization Matters: Skipping L2 normalization will break the cosine similarity logic, so don't skip that step.
- Task Flexibility: If you're using this for classification (e.g., "is this answer a match?"), add a dense output layer on top of the cosine distance and use binary crossentropy loss instead.
内容的提问来源于stack exchange,提问作者Slowpoke

