You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Triplet Loss网络训练收敛至Constant Embeddings问题求助

我正在搭建的Siamese模型抽象结构如下图所示:
Siamese模型结构

在DeepFashion数据集上训练该模型时,损失随训练推进逐步上升直至达到恒定值,此时模型会输出恒定嵌入向量(Constant Embeddings)。
训练损失变化曲线

为验证所有嵌入是否完全相同,我计算了所有Anchor、Positive、Negative三元组的距离,验证代码如下:

is_constant = lambda a, b: tf.math.reduce_sum((a - b)) == 0

for batch in (dataset["val"].take(1)): #siamese_model.validate_embedding
  anchors, positives, negatives = siamese_network(a)
  for a, p, n in (zip(anchors, positives, negative)):
    all_const = all([is_constant(a, p), is_constant(a, n), is_constant(p, n)])
  
    print(a)
    print(p)
    print(n)
    print("All Embds. Constant", all_const)
    print("*"*72)
  break

运行代码后得到重复输出(重复次数等于批次大小)如下:

tf.Tensor(
[-0.03565513 -0.06550519  0.01807223 ... -0.0639673  -0.01235839
  0.04788744], shape=(2048,), dtype=float32)
tf.Tensor(
[-0.03565513 -0.06550519  0.01807223 ... -0.0639673  -0.01235839
  0.04788744], shape=(2048,), dtype=float32)
tf.Tensor(
[-0.03565513 -0.06550519  0.01807223 ... -0.0639673  -0.01235839
  0.04788744], shape=(2048,), dtype=float32)
All Embds. Constant True
************************************************************************

目前我尚未找到该问题的解决方案。

ResNet骨干网络

我的ResNet骨干网络实现代码如下:

class ResnetIdentityBlock(tf.keras.Model):
    def __init__(self, kernel_size, filters) -> object:
        super(ResnetIdentityBlock, self).__init__(name='')
        filters1, filters2, filters3 = filters

        self.conv2a = tf.keras.layers.Conv2D(filters1, (1, 1), name="conv2a")
        self.bn2a = tf.keras.layers.BatchNormalization(name="bn2a")

        self.conv2b = tf.keras.layers.Conv2D(filters2, kernel_size, padding='same', name="conv2b")
        self.bn2b = tf.keras.layers.BatchNormalization(name="bn2b")

        self.conv2c = tf.keras.layers.Conv2D(filters3, (1, 1), name="conv2c")
        self.bn2c = tf.keras.layers.BatchNormalization(name="bn2c")

    def call(self, input_tensor, training=False):
        x = self.conv2a(input_tensor)
        x = self.bn2a(x, training=training)
        x = tf.nn.relu(x)

        x = self.conv2b(x)
        x = self.bn2b(x, training=training)
        x = tf.nn.relu(x)

        x = self.conv2c(x)
        x = self.bn2c(x, training=training)

        x += input_tensor
        return tf.nn.relu(x)
input_shape = (224, 224)
weights = "imagenet"  

R = resnet.ResNet50(
    weights=weights, input_shape=input_shape + (3,), include_top=False
)

embedding_model = Sequential([
    R,
    tf.keras.layers.Conv2D(256, (3, 3), activation='relu',
        kernel_initializer='he_uniform', padding='same',
        kernel_regularizer=l2(2e-4)),
    ResnetIdentityBlock(3, [64, 64, 256]),
    tf.keras.layers.AveragePooling2D(pool_size=(2, 2), strides=(1, 1), padding='same'),
    tf.keras.layers.Flatten(),
    tf.keras.layers.Dense(embedding_dim, activation=None, use_bias=False)
])

上述骨干网络被封装为SiameseNetwork,会为所有输入预测Embedding并计算数据点间的距离(Anchor与Positive、Anchor与Negative),封装代码如下:

anchor_input = layers.Input(name="anchor", shape=input_shape +(3,))
positive_input = layers.Input(name="positive", shape=input_shape +(3,))
negative_input = layers.Input(name="negative", shape=input_shape +(3,))

anchor_encoding   = back_bone(resnet.preprocess_input(anchor_input))
positive_encoding = back_bone(resnet.preprocess_input(positive_input))
negative_encoding = back_bone(resnet.preprocess_input(negative_input))

loss_output = TripletLoss(alpha=1.0)(anchor_encoding, positive_encoding, negative_encoding)

siamese_network = Model(
    inputs=[anchor_input, positive_input, negative_input], 
    outputs=loss_output
)

输出层负责计算Triplet-Loss,我自定义的Triplet-Loss层实现如下:

@tf.function
def euclidean_distance(x, y):
    l2_distance = tf.math.square(tf.subtract(x, y))
    return tf.math.reduce_sum(l2_distance, axis=-1)
  
class TripletLoss(layers.Layer):
    def __init__(self, alpha, **kwargs):
        super(TripletLoss, self).__init__(**kwargs)
        self.alpha = alpha # im using Alpha = 1.0
        self.distance_layer = TripletDistance()

    def call(self, anchor, positive, negative):
        ap_distance = euclidean_distance(anchor, positive)
        an_distance = euclidean_distance(anchor, negative)

        loss = ap_distance - an_distance
        loss = tf.maximum(loss + self.alpha, 0.0)

        return loss

优化器我使用的是学习率为1e-4的Adam优化器。

请问我的实现是否存在本质错误?


内容的提问来源于stack exchange,提问作者Barney Stinson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 19:54:08