You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在两层神经网络中正确实现Center Loss及其导数

Center Loss实现问题排查与修正

我搭好了基础两层神经网络,但实现Center Loss后达不到原论文的聚类效果——损失曲线正常,但特征没有类内聚集。参考了TensorFlow的L2损失方案仍无进展,相关代码如下:

import numpy as np
import pandas as pd
import tensorflow as tf
import sklearn as sk
from sklearn import preprocessing
from tensorflow.keras.datasets import mnist
import matplotlib.pyplot as plt

(x_train,y_train),(x_test,y_test)=mnist.load_data()

n_train = 100
x_train = x_train[0:n_train].reshape(-1,784)/255
y_train0 = y_train[0:n_train].reshape(-1,)
y_train1 = pd.get_dummies(y_train0)
y_train = np.array( y_train1.astype(int) )


nx = 784 #input size (nx,1)
n1 = 20   #neurons in 1st layer
n_class = 10  #neurons in 2nd layer


lambda0 = 0.05  #center loss parameter
alpha =  0.001  #模型学习率
center_alpha = 0.5  #Center更新学习率(原论文推荐0.5)

#layer weights
#The network is simple 
#Input, n1 neurons with sigmoid activation, 10 neurons and softmax output
W1 = 0.001*np.random.rand( n1, nx )
W2 = 0.001*np.random.rand( n_class, n1 )
b2 = 0.001*np.random.rand( n_class,1  )
b1 = 0.001*np.random.rand( n1,1  )
centers = np.random.rand(n_class, n1)  #Center维度和A1一致(n_class, n1)


def softmax(x):
    # 标准softmax实现,避免数值溢出
    exp_x = np.exp(x - np.max(x, axis=0, keepdims=True))
    return exp_x / np.sum(exp_x, axis=0, keepdims=True)

def sigmoid(x):
    y = 1/(1+ np.exp(-x) )
    return y


def test():
    count = 0
    feature_list = []
    label_list = []
    for t3 in range(0,100):
        #forward propagation #Layer 1
        Z1 = np.dot( W1 , x_train[t3]).reshape(-1,1) + b1
        A1 = sigmoid( Z1 ).reshape(-1,1)         
        #forward propagation #Layer 2
        Z2 = np.dot( W2, A1)  + b2
        Output = softmax(Z2)
        
        arg_max =  np.argmax(Output)
        if arg_max == np.argmax(y_train[t3:t3+1] ):
            count = count +1
        feature_list.append(A1.flatten())
        label_list.append(y_train0[t3])
    
    print(count/100)
    return np.array(feature_list), np.array(label_list)


iter = 1500
sets = 100
loss = np.zeros(( iter,1))

print_counter = 0
for t1 in range (0,iter):
    total_loss = 0.0
    feature_epoch = []
    label_epoch = []
    for t2 in range(0,sets):
        #forward propagation #Layer 1 & 2
        x = x_train[t2:t2+1].T  # shape (784,1)
        Z1 = np.dot( W1 , x ) + b1  # shape (20,1)
        A1 = sigmoid( Z1 )  # shape (20,1)
        Z2 = np.dot( W2, A1)  + b2  # shape (10,1)
        Output = softmax(Z2)  # shape (10,1)
        
        # 获取当前样本的类别索引
        y_idx = np.argmax(y_train[t2:t2+1])
        
        # 计算交叉熵损失
        cross_entropy = -np.sum( y_train[t2:t2+1].T * np.log(Output + 1e-8) )
        # 计算Center Loss(L2平方损失)
        center_loss = 0.5 * np.sum( (A1 - centers[y_idx].reshape(-1,1)) **2 )
        # 总损失
        total_loss_step = cross_entropy + lambda0 * center_loss
        total_loss += total_loss_step
        
        # 反向传播:先算总损失对各变量的梯度
        # 对Z2的梯度(交叉熵部分)
        dZ2_cross = Output - y_train[t2:t2+1].T
        # 对A1的梯度(Center Loss部分)
        dA1_center = A1 - centers[y_idx].reshape(-1,1)
        
        # 总梯度:Z2的梯度仅来自交叉熵,A1的梯度来自交叉熵+Center Loss
        dZ2 = dZ2_cross
        dA1 = np.dot(W2.T, dZ2) + lambda0 * dA1_center
        
        # 计算各权重和偏置的梯度
        dW2 = np.dot(dZ2, A1.T)
        db2 = dZ2
        dZ1 = dA1 * A1 * (1 - A1)  # sigmoid的导数
        dW1 = np.dot(dZ1, x.T)
        db1 = dZ1
        
        # 更新模型参数
        W1 = W1 - alpha * dW1
        W2 = W2 - alpha * dW2
        b1 = b1 - alpha * db1
        b2 = b2 - alpha * db2
        
        # 更新Center(原论文公式)
        centers[y_idx] = centers[y_idx] + center_alpha * (A1.flatten() - centers[y_idx])
        
        feature_epoch.append(A1.flatten())
        label_epoch.append(y_idx)
    
    loss[t1] = total_loss / sets
    
    print_counter = print_counter + 1
    if print_counter > 100:
        print(f"迭代{t1},损失:{loss[t1][0]:.4f}")
        print_counter = 1 

# 测试并获取特征
features, labels = test()

plt.plot(loss )
plt.xlabel('迭代次数')
plt.ylabel('总损失')
plt.show()

# 用PCA降维可视化特征(因为A1是20维,取前2维)
from sklearn.decomposition import PCA
pca = PCA(n_components=2)
features_2d = pca.fit_transform(features)

plt.figure(figsize=(8,6))
for i in range(n_class):
    mask = labels == i
    plt.scatter(features_2d[mask,0], features_2d[mask,1], label=f"类别{i}", alpha=0.7)
plt.legend()
plt.xlabel('PCA维度1')
plt.ylabel('PCA维度2')
plt.title('Center Loss优化后的特征聚类')
plt.show()

关键修正说明

  • 损失函数修正:改用原论文的L2平方损失,将Center Loss加入总损失,让模型同时优化分类和聚类目标。
  • Center更新逻辑:按照原论文的增量更新方式,而非直接取均值,保证Center的更新和模型训练同步。
  • 特征选择:使用倒数第二层的A1作为聚类特征,这层特征更具代表性,适合做类内聚集优化。
  • 数值稳定性:优化了softmax实现,避免数值溢出;损失计算加入1e-8防止log(0)错误。
  • 可视化改进:用PCA将20维特征降维到2维,更直观展示类内聚集效果。

内容的提问来源于stack exchange,提问作者gingerorange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 20:24:54