如何计算各类别中心间距并最大化MNIST类别中心间距离?
Hey there! Let's fix this issue. The problem with using tf.add(centres, margin) is that it's just a static offset—it doesn't integrate with your training process to learn to pull the centers apart. To make this work, you need a differentiable loss function that tells the optimizer to push different class centers as far away from each other as possible.
Here's a step-by-step implementation:
1. Ensure Your Centers Are Trainable
First, double-check that your centers variable is set to be trainable (this is default, but it's easy to miss):
num_classes = 10 len_features = 2 centers = tf.get_variable('centers', [num_classes, len_features], dtype=tf.float32, initializer=tf.initializers.random_normal(), # Or use your existing initialization logic trainable=True)
2. Calculate Pairwise Distances Between Centers
Use TensorFlow's broadcasting to compute the squared distance between every pair of centers (squared distance avoids the expensive square root operation while preserving the same optimization direction):
# Expand dimensions to enable broadcasting across all center pairs centers_expanded = tf.expand_dims(centers, axis=0) # Shape: (1, 10, 2) centers_expanded_t = tf.expand_dims(centers, axis=1) # Shape: (10, 1, 2) # Compute squared pairwise distances: final shape (10, 10) pairwise_sq_distances = tf.reduce_sum(tf.square(centers_expanded - centers_expanded_t), axis=-1)
3. Create a Mask for Non-Same-Class Pairs
We don't care about the distance between a center and itself (which is 0), so create a mask to ignore those diagonal elements:
# Mask where diagonal (same class) is 0, all other pairs are 1 mask = 1 - tf.eye(num_classes, dtype=tf.float32)
4. Build the "Pull Apart" Loss
There are two common effective ways to define this loss:
Option A: Hinge Loss (Enforce Minimum Margin)
If you want all different class centers to be at least a certain margin apart, use a hinge loss. This penalizes pairs that are closer than your desired threshold:
margin = 2.0 # Adjust this based on your 2D feature scale # Compute hinge loss: penalize pairs where distance < margin pull_apart_loss = tf.reduce_sum(tf.maximum(margin - tf.sqrt(pairwise_sq_distances), 0) * mask)
Option B: Negative Distance Sum (Maximize All Distances)
If you just want to maximize all pairwise distances as much as possible, use the negative sum of squared distances. Minimizing this loss directly translates to maximizing the actual distances between centers:
pull_apart_loss = -tf.reduce_sum(pairwise_sq_distances * mask)
5. Combine with Your Classification Loss
Add this loss to your existing classification loss (e.g., cross-entropy) with a weighting parameter to balance the two objectives:
lambda_param = 0.1 # Tune this value based on your training results (start small) classification_loss = ... # Your existing classification loss (e.g., sparse categorical crossentropy) total_loss = classification_loss + lambda_param * pull_apart_loss
6. Optimize the Total Loss
Use your favorite optimizer to minimize the total loss—this will update both your model weights and the class centers to improve classification and pull centers apart:
optimizer = tf.train.AdamOptimizer(learning_rate=0.001) train_op = optimizer.minimize(total_loss)
Why Your Original Approach Failed
tf.add(centres, margin) just shifts all centers by a fixed value—it doesn't consider the relative positions of the centers, and it's not connected to the training loop. The optimizer can't learn from this because there's no gradient signal telling it how to adjust centers to maximize pairwise distances.
内容的提问来源于stack exchange,提问作者Kamran Janjua

