Siamese网络训练无效问题排查求助
Hey there, let’s dig into why your siamese network is stuck at 50% accuracy—it’s essentially guessing randomly, so we need to fix the core mismatches and gaps in your setup step by step:
1. Critical Mismatch: Labels vs. Loss Function
This is probably the biggest issue:
- Contrastive loss expects scalar labels (0 or 1), not one-hot encoded ones from
to_categorical. The loss logic is:y_true=1means the two images are a matching pair,y_true=0means they’re not. Using one-hot labels breaks how the loss calculates similarity/dissimilarity. - When you tried
categorical_crossentropy, your model output is still an Euclidean distance (a continuous value), not a class probability distribution. This loss function can’t interpret distance values as classification outputs.
Fix this first:
# Revert your labels to a 1D array of 0s and 1s (no to_categorical) y_train = ... # Shape should be (num_samples,) instead of (num_samples, 2)
2. Invalid Accuracy Metric
The accuracy metric you’re using is meaningless here! Keras’s default accuracy compares model outputs directly to labels, but your model outputs a distance value, not a 0/1 prediction. You need a custom metric for contrastive learning:
def contrastive_accuracy(y_true, y_pred): # Set a threshold (adjust based on your distance distribution) threshold = 0.5 # Predict "match" if distance < threshold, "no match" otherwise predictions = K.cast(y_pred < threshold, y_true.dtype) return K.mean(K.equal(y_true, predictions))
Then use it when compiling:
model.compile(loss=contrastive_loss, optimizer='RMSprop', metrics=[contrastive_accuracy])
3. Image Data Issues
- No normalization: Your input is flattened image pixels (9216 = 96x96), which are likely in 0-255 range. This large scale can destabilize training. Normalize to 0-1:
tr_pair1_reshaped = tr_pair1_reshaped / 255.0 tr_pair2_reshaped = tr_pair2_reshaped / 255.0 - Verify pair correctness: Double-check that your
tr_pair1andtr_pair2actually correspond to the labels. For example, wheny_train[i] = 1, aretr_pair1[i]andtr_pair2[i]truly the same class? A mix-up here would make learning impossible. - Check class balance: If your 40k pairs are perfectly split 50/50 between matches and non-matches, random guessing gives 50% accuracy. But confirm the distribution isn’t skewed (though 50% suggests it’s balanced, but better to be sure).
4. Model Structure is Suboptimal for Images
A fully connected network is terrible at capturing spatial patterns in images—you need a CNN instead! Replace your create_base_network with a convolutional feature extractor:
def create_base_network(input_shape): seq = Sequential() # Adjust input_shape to match your images (e.g., (96,96,1) for grayscale, (96,96,3) for RGB) seq.add(Conv2D(32, (3,3), activation='relu', input_shape=input_shape)) seq.add(MaxPooling2D((2,2))) seq.add(Dropout(0.3)) seq.add(Conv2D(64, (3,3), activation='relu')) seq.add(MaxPooling2D((2,2))) seq.add(Dropout(0.3)) seq.add(Conv2D(128, (3,3), activation='relu')) seq.add(MaxPooling2D((2,2))) seq.add(Dropout(0.3)) seq.add(Flatten()) seq.add(Dense(128, activation='relu')) return seq # Update input definitions to use 2D image shape instead of flattened vector input_shape = (96,96,1) # Adjust based on your image format input_a = Input(shape=input_shape) input_b = Input(shape=input_shape) base_network = create_base_network(input_shape) processed_a = base_network(input_a) processed_b = base_network(input_b)
5. Training Hyperparameters
- Too few epochs: 3 epochs is nowhere near enough for 40k samples. Try at least 30-50 epochs and monitor the loss/metric to see if it improves over time.
- Learning rate: If you didn’t normalize data, default learning rates (0.001 for Adam, 0.001 for RMSprop) might be too large. Try reducing to 0.0001 and see if training stabilizes.
- Dropout rate: 0.1 is very low for fully connected layers. If you stick with the original dense network, bump dropout to 0.3-0.5 to prevent overfitting.
Bonus: Alternative Setup for Binary Classification
If you prefer using standard classification metrics, modify your model to output a probability instead of distance:
# Add a sigmoid layer to convert distance to a 0-1 match probability distance = Lambda(euclidean_distance, output_shape=eucl_dist_output_shape)([processed_a, processed_b]) output = Dense(1, activation='sigmoid')(distance) model = Model(inputs=[input_a, input_b], outputs=output) # Compile with binary crossentropy (uses scalar 0/1 labels) model.compile(loss='binary_crossentropy', optimizer='Adam', metrics=['accuracy'])
Follow these steps in order—fixing the label/loss mismatch and adding a CNN should immediately move your accuracy away from 50%.
内容的提问来源于stack exchange,提问作者Arka Mallick

