You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow自定义CosFace层调用异常及训练精度异常排查

Fixing Your CosFace Layer Issues in TensorFlow

Hey, let's break down what's going wrong here and fix it step by step—you've got two main problems, both rooted in small but critical code mistakes.

First, Let's Diagnose the Issues

  1. Impossible 1.0 Accuracy & Broken Loss Logic
    Your CosFace layer's core calculation is inverted. Let's look at your code:

    theta = tf.math.acos(K.clip(cosine_sim, -1.0 + K.epsilon(), 1.0 - K.epsilon()))
    final_theta = tf.where(tf.cast(one_hot_labels, dtype=tf.bool), tf.math.cos(theta) - self._m, tf.math.cos(theta), name='final_theta')
    output = tf.math.cos(final_theta, name='cosine_sim_with_margin')
    

    CosFace's formula is s * (cosθ - m) for positive classes, but you're taking cos(cosθ - m)—that's a nonsensical transformation that blows out your logits, making the softmax output extreme (hence the fake 1.0 accuracy). You don't need to convert cosine similarity to radians and back; just modify the cosine value directly.

  2. Print Statements Disappearing After Epoch 1
    Python's print only runs once when TensorFlow builds the computation graph. During training (in graph mode), it won't execute again. You need to use tf.print instead—it's embedded into the graph and runs every time the layer is called.

  3. Incorrect Label Casting
    Using int(label) on a Tensor is invalid; you have to use TensorFlow's tf.cast to convert the label tensor to integer type.

  4. Incomplete Layer Build
    Your build method doesn't call super().build(input_shape), which leaves the layer's internal state uninitialized properly.

Fixed CosFace Layer Code

Here's the corrected implementation:

import math
import numpy as np
import tensorflow as tf
from tensorflow import keras
import tensorflow.keras.backend as K
from tensorflow.keras.layers import Layer
from tensorflow.python.keras.utils import tf_utils

def _resolve_training(layer, training):
    if training is None:
        training = K.learning_phase()
    if isinstance(training, int):
        training = bool(training)
    if not layer.trainable:
        training = False
    return tf_utils.constant_value(training)

class CosFace(keras.layers.Layer):
    """ Implementation of CosFace layer.
    Reference: https://arxiv.org/abs/1801.09414
    Arguments:
        num_classes: number of classes to classify
        s: scale factor
        m: margin
        regularizer: weights regularizer
    """
    def __init__(self, num_classes, s=30.0, m=0.35, regularizer=None, name='cosface', **kwargs):
        super().__init__(name=name, **kwargs)
        self._n_classes = num_classes
        self._s = float(s)
        self._m = float(m)
        self._regularizer = regularizer

    def build(self, input_shape):
        embedding_shape, label_shape = input_shape
        self._w = self.add_weight(shape=(embedding_shape[-1], self._n_classes),
                                  initializer='glorot_uniform',
                                  trainable=True,
                                  regularizer=self._regularizer)
        # Don't forget to call super's build to finalize the layer
        super().build(input_shape)

    def call(self, inputs, training=None):
        """ During training, requires 2 inputs: embedding (after backbone+pool+dense), and ground truth labels.
        The labels should be sparse (and use sparse_categorical_crossentropy as loss).
        """
        # Use tf.print instead of Python print for graph-mode execution
        tf.print('calling CosFace Layer...')
        embedding, label = inputs
        # Correct way to cast and reshape labels
        label = tf.reshape(tf.cast(label, tf.int32), [-1], name='label_shape_correction')
        
        # Normalize embeddings and weights
        x = tf.nn.l2_normalize(embedding, axis=1, name='normalize_prelogits')
        w = tf.nn.l2_normalize(self._w, axis=0, name='normalize_weights')
        cosine_sim = tf.matmul(x, w, name='cosine_similarity')

        training = _resolve_training(self, training)
        if not training:
            return self._s * cosine_sim
        else:
            one_hot_labels = tf.one_hot(label, depth=self._n_classes, name='one_hot_labels')
            # Correct CosFace logic: subtract margin only from positive class cosine values
            cosine_sim_with_margin = tf.where(
                tf.cast(one_hot_labels, dtype=tf.bool),
                cosine_sim - self._m,
                cosine_sim,
                name='cosine_sim_with_margin'
            )
            tf.print('Cosine sim with margin shape:', cosine_sim_with_margin.shape)
            return self._s * cosine_sim_with_margin

Additional Notes for Your Model

  1. Verify Training Inputs
    Ensure y_train is an integer array (not one-hot encoded) since you're using sparse_categorical_crossentropy.
  2. Check Optimizer Definition
    Your code references opt but doesn't show its definition—make sure it's properly initialized (e.g., opt = keras.optimizers.Adam()).

Why This Fixes Your Problems

  • The corrected CosFace logic applies the margin directly to the cosine similarity, which aligns with the original paper and produces valid logits. This will fix the fake 1.0 accuracy issue, and you'll see realistic accuracy values during training.
  • tf.print will now log output every time the layer is called, so you'll see messages in every epoch.
  • The layer is fully initialized with super().build(input_shape), preventing any hidden state issues.

内容的提问来源于stack exchange,提问作者Lorenzo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 14:47:33