如何在两层神经网络中正确实现Center Loss及其导数
Center Loss实现问题排查与修正
我搭好了基础两层神经网络,但实现Center Loss后达不到原论文的聚类效果——损失曲线正常,但特征没有类内聚集。参考了TensorFlow的L2损失方案仍无进展,相关代码如下:
import numpy as np import pandas as pd import tensorflow as tf import sklearn as sk from sklearn import preprocessing from tensorflow.keras.datasets import mnist import matplotlib.pyplot as plt (x_train,y_train),(x_test,y_test)=mnist.load_data() n_train = 100 x_train = x_train[0:n_train].reshape(-1,784)/255 y_train0 = y_train[0:n_train].reshape(-1,) y_train1 = pd.get_dummies(y_train0) y_train = np.array( y_train1.astype(int) ) nx = 784 #input size (nx,1) n1 = 20 #neurons in 1st layer n_class = 10 #neurons in 2nd layer lambda0 = 0.05 #center loss parameter alpha = 0.001 #模型学习率 center_alpha = 0.5 #Center更新学习率(原论文推荐0.5) #layer weights #The network is simple #Input, n1 neurons with sigmoid activation, 10 neurons and softmax output W1 = 0.001*np.random.rand( n1, nx ) W2 = 0.001*np.random.rand( n_class, n1 ) b2 = 0.001*np.random.rand( n_class,1 ) b1 = 0.001*np.random.rand( n1,1 ) centers = np.random.rand(n_class, n1) #Center维度和A1一致(n_class, n1) def softmax(x): # 标准softmax实现,避免数值溢出 exp_x = np.exp(x - np.max(x, axis=0, keepdims=True)) return exp_x / np.sum(exp_x, axis=0, keepdims=True) def sigmoid(x): y = 1/(1+ np.exp(-x) ) return y def test(): count = 0 feature_list = [] label_list = [] for t3 in range(0,100): #forward propagation #Layer 1 Z1 = np.dot( W1 , x_train[t3]).reshape(-1,1) + b1 A1 = sigmoid( Z1 ).reshape(-1,1) #forward propagation #Layer 2 Z2 = np.dot( W2, A1) + b2 Output = softmax(Z2) arg_max = np.argmax(Output) if arg_max == np.argmax(y_train[t3:t3+1] ): count = count +1 feature_list.append(A1.flatten()) label_list.append(y_train0[t3]) print(count/100) return np.array(feature_list), np.array(label_list) iter = 1500 sets = 100 loss = np.zeros(( iter,1)) print_counter = 0 for t1 in range (0,iter): total_loss = 0.0 feature_epoch = [] label_epoch = [] for t2 in range(0,sets): #forward propagation #Layer 1 & 2 x = x_train[t2:t2+1].T # shape (784,1) Z1 = np.dot( W1 , x ) + b1 # shape (20,1) A1 = sigmoid( Z1 ) # shape (20,1) Z2 = np.dot( W2, A1) + b2 # shape (10,1) Output = softmax(Z2) # shape (10,1) # 获取当前样本的类别索引 y_idx = np.argmax(y_train[t2:t2+1]) # 计算交叉熵损失 cross_entropy = -np.sum( y_train[t2:t2+1].T * np.log(Output + 1e-8) ) # 计算Center Loss(L2平方损失) center_loss = 0.5 * np.sum( (A1 - centers[y_idx].reshape(-1,1)) **2 ) # 总损失 total_loss_step = cross_entropy + lambda0 * center_loss total_loss += total_loss_step # 反向传播:先算总损失对各变量的梯度 # 对Z2的梯度(交叉熵部分) dZ2_cross = Output - y_train[t2:t2+1].T # 对A1的梯度(Center Loss部分) dA1_center = A1 - centers[y_idx].reshape(-1,1) # 总梯度:Z2的梯度仅来自交叉熵,A1的梯度来自交叉熵+Center Loss dZ2 = dZ2_cross dA1 = np.dot(W2.T, dZ2) + lambda0 * dA1_center # 计算各权重和偏置的梯度 dW2 = np.dot(dZ2, A1.T) db2 = dZ2 dZ1 = dA1 * A1 * (1 - A1) # sigmoid的导数 dW1 = np.dot(dZ1, x.T) db1 = dZ1 # 更新模型参数 W1 = W1 - alpha * dW1 W2 = W2 - alpha * dW2 b1 = b1 - alpha * db1 b2 = b2 - alpha * db2 # 更新Center(原论文公式) centers[y_idx] = centers[y_idx] + center_alpha * (A1.flatten() - centers[y_idx]) feature_epoch.append(A1.flatten()) label_epoch.append(y_idx) loss[t1] = total_loss / sets print_counter = print_counter + 1 if print_counter > 100: print(f"迭代{t1},损失:{loss[t1][0]:.4f}") print_counter = 1 # 测试并获取特征 features, labels = test() plt.plot(loss ) plt.xlabel('迭代次数') plt.ylabel('总损失') plt.show() # 用PCA降维可视化特征(因为A1是20维,取前2维) from sklearn.decomposition import PCA pca = PCA(n_components=2) features_2d = pca.fit_transform(features) plt.figure(figsize=(8,6)) for i in range(n_class): mask = labels == i plt.scatter(features_2d[mask,0], features_2d[mask,1], label=f"类别{i}", alpha=0.7) plt.legend() plt.xlabel('PCA维度1') plt.ylabel('PCA维度2') plt.title('Center Loss优化后的特征聚类') plt.show()
关键修正说明
- 损失函数修正:改用原论文的L2平方损失,将Center Loss加入总损失,让模型同时优化分类和聚类目标。
- Center更新逻辑:按照原论文的增量更新方式,而非直接取均值,保证Center的更新和模型训练同步。
- 特征选择:使用倒数第二层的
A1作为聚类特征,这层特征更具代表性,适合做类内聚集优化。 - 数值稳定性:优化了softmax实现,避免数值溢出;损失计算加入
1e-8防止log(0)错误。 - 可视化改进:用PCA将20维特征降维到2维,更直观展示类内聚集效果。
内容的提问来源于stack exchange,提问作者gingerorange
相关产品推荐
相关产品推荐

