图像二分类用软标签准确率骤降,损失函数及问题排查咨询
问题
我们在TensorFlow中用CNN做标签为0和1的图像二分类,实际场景里图像标签是0到1之间的概率值(非0/1独热标签):[0,0.5)区间记为0,[0.5,1.0]区间记为1。现在想验证用0-1软标签替代独热标签能不能提升分类性能,但基于CIFAR10数据集0、1类的示例代码,未添加软标签时准确率约98%,添加后准确率降至约48%。
想问:
- 是否需要修改自定义损失函数
BinaryCrossEntropy_custom?还是存在其他问题? - 参考资料提到用logits能解决问题,我们理解soft_labels参数,但这个示例里logits参数该传入什么值?
附示例代码:
from tensorflow.keras import datasets, layers, models import matplotlib.pyplot as plt import numpy as np from tensorflow.keras import optimizers from tensorflow.keras.optimizers import Adam from tensorflow.keras.applications import vgg16 from tensorflow.keras.models import Model from tensorflow.keras import backend as K # Rewrite the binary cross entropy function. We will modify this function to return a loss that fits the soft label later. def BinaryCrossEntropy_custom(y_true, y_pred): y_pred = K.clip(y_pred, K.epsilon(), 1 - K.epsilon()) term_0 = (1 - y_true) * K.log(1 - y_pred + K.epsilon()) term_1 = y_true * K.log(y_pred + K.epsilon()) return -K.mean(term_0 + term_1, axis=0) (train_images, train_labels), (test_images, test_labels) = datasets.cifar10.load_data() # Normalize pixel values to be between 0 and 1 train_images, test_images = train_images / 255.0, test_images / 255.0 # Only data with labels 0 and 1 are used. train_ind01 = np.where((train_labels == 0) | (train_labels == 1))[0] test_ind01 = np.where((test_labels == 0) | (test_labels == 1))[0] train_images = train_images[train_ind01, :, :, :] test_images = test_images[test_ind01, :, :, :] train_labels = train_labels[train_ind01, :] test_labels = test_labels[test_ind01, :] train_labels = np.array(train_labels).astype('float64') test_labels = np.array(test_labels).astype('float64') # making soft labels part start # Samples with label 0 are replaced with labels in the range [0,0.2], # and samples with label 1 are replaced by labels in the range [0.8, 1.0]. sampl_train = np.random.uniform(low=-0.2, high=0.2, size=train_labels.shape) sampl_test = np.random.uniform(low=-0.2, high=0.2, size=test_labels.shape) train_labels = train_labels + sampl_train test_labels = test_labels + sampl_test train_labels = np.clip(train_labels, 0.0, 1.0) test_labels = np.clip(test_labels, 0.0, 1.0) # making soft labels part end vgg = vgg16.VGG16(include_top=False, weights='imagenet', input_shape=(32, 32, 3)) output = vgg.layers[-1].output output = layers.Flatten()(output) output = layers.Dense(512, activation='relu')(output) output = layers.Dropout(0.2)(output) output = layers.Dense(256, activation='relu')(output) output = layers.Dropout(0.2)(output) predictions = layers.Dense(units=1, activation="sigmoid")(output) model = Model(inputs=vgg.input, outputs=predictions) model.compile(optimizer=Adam(learning_rate=.0001), loss=BinaryCrossEntropy_custom, metrics=['accuracy']) history = model.fit(train_images, train_labels, epochs=100, validation_data=(test_images, test_labels)) plt.plot(history.history['accuracy'], label='accuracy') plt.plot(history.history['val_accuracy'], label='val_accuracy') plt.xlabel('Epoch') plt.ylabel('Accuracy') plt.ylim([0.5, 1]) plt.legend(loc='lower right') test_loss, test_acc = model.evaluate(test_images, test_labels, verbose=2) print(test_acc)
添加软标签后的控制台输出:
Epoch 1/100 2022-09-16 15:29:29.136931: I tensorflow/stream_executor/cuda/cuda_dnn.cc:366] Loaded cuDNN version 8101 313/313 [==============================] - 17s 42ms/step - loss: 0.2951 - accuracy: 0.4779 - val_loss: 0.2775 - val_accuracy: 0.4650 Epoch 2/100 313/313 [==============================] - 12s 38ms/step - loss: 0.2419 - accuracy: 0.4931 - val_loss: 0.2488 - val_accuracy: 0.4695 Epoch 3/100 313/313 [==============================] - 12s 39ms/step - loss: 0.2290 - accuracy: 0.4978 - val_loss: 0.2424 - val_accuracy: 0.4740 Epoch 4/100 313/313 [==============================] - 12s 39ms/step - loss: 0.2161 - accuracy: 0.5002 - val_loss: 0.2404 - val_accuracy: 0.4765 Epoch 5/100 313/313 [==============================] - 12s 39ms/step - loss: 0.2139 - accuracy: 0.5007 - val_loss: 0.2620 - val_accuracy: 0.4730 Epoch 6/100 313/313 [==============================] - 12s 38ms/step - loss: 0.2118 - accuracy: 0.5023 - val_loss: 0.2480 - val_accuracy: 0.4745 Epoch 7/100 313/313 [==============================] - 12s 38ms/step - loss: 0.2097 - accuracy: 0.5019 - val_loss: 0.2350 - val_accuracy: 0.4775 Epoch 8/100 313/313 [==============================] - 12s 39ms/step - loss: 0.2098 - accuracy: 0.5024 - val_loss: 0.2289 - val_accuracy: 0.4780 Epoch 9/100 313/313 [==============================] - 12s 38ms/step - loss: 0.2034 - accuracy: 0.5039 - val_loss: 0.2364 - val_accuracy: 0.4780 Epoch 10/100 313/313 [==============================] - 12s 39ms/step - loss: 0.2025 - accuracy: 0.5040 - val_loss: 0.2481 - val_accuracy: 0.4720
核心问题分析
代码里有两个关键问题导致准确率暴跌:
- 准确率指标逻辑不匹配软标签:Keras默认的
accuracy指标是将预测值和真实值做硬匹配(比如预测>0.5算1,否则算0,再和真实标签的0/1对比),但你现在的真实标签是软标签(比如0.1或0.9),默认指标会把所有软标签当成非0即1的硬标签对比,逻辑完全混乱,准确率接近随机猜测。 - 自定义损失函数本身没问题,但可替换为官方实现:你写的
BinaryCrossEntropy_custom和TensorFlow官方的BinaryCrossentropy(from_logits=False时)逻辑一致,本身兼容软标签,无需修改,但官方实现更稳定。
具体修复方案
1. 替换准确率指标,适配软标签
不要用默认的accuracy,改成自定义逻辑:先把软标签转成硬标签(y_true >=0.5记为1,否则0),再和预测值的硬分类结果对比。
2. 关于logits的使用
如果要用logits模式,需做两处修改:
- 模型最后一层去掉
sigmoid激活,直接输出原始logit值(未经过sigmoid压缩的实数) - 损失函数使用
BinaryCrossentropy(from_logits=True),此时损失函数内部会自动对logit做sigmoid转换,再计算交叉熵,数值稳定性更好,也兼容软标签。
3. 优化软标签生成逻辑(非必须,更直观)
原代码生成软标签的逻辑没问题,但可以直接生成指定区间的随机数,避免裁剪操作:
train_labels = np.where(train_labels == 0, np.random.uniform(0, 0.2, size=train_labels.shape), np.random.uniform(0.8, 1.0, size=train_labels.shape)) test_labels = np.where(test_labels == 0, np.random.uniform(0, 0.2, size=test_labels.shape), np.random.uniform(0.8, 1.0, size=test_labels.shape))
修改后的完整代码
from tensorflow.keras import datasets, layers, models import matplotlib.pyplot as plt import numpy as np from tensorflow.keras.optimizers import Adam from tensorflow.keras.applications import vgg16 from tensorflow.keras.models import Model from tensorflow.keras import backend as K from tensorflow.keras.losses import BinaryCrossentropy # 自定义适配软标签的准确率指标 def soft_label_accuracy(y_true, y_pred): # 软标签转硬标签:>=0.5为1,否则为0 y_true_hard = K.cast(y_true >= 0.5, K.floatx()) # 预测值转硬标签:>0.5为1,否则为0 y_pred_hard = K.cast(y_pred >= 0.5, K.floatx()) return K.mean(K.equal(y_true_hard, y_pred_hard), axis=-1) (train_images, train_labels), (test_images, test_labels) = datasets.cifar10.load_data() # 归一化像素值 train_images, test_images = train_images / 255.0, test_images / 255.0 # 只保留标签0和1的数据 train_ind01 = np.where((train_labels == 0) | (train_labels == 1))[0] test_ind01 = np.where((test_labels == 0) | (test_labels == 1))[0] train_images = train_images[train_ind01, :, :, :] test_images = test_images[test_ind01, :, :, :] train_labels = train_labels[train_ind01, :] test_labels = test_labels[test_ind01, :] train_labels = np.array(train_labels).astype('float64') test_labels = np.array(test_labels).astype('float64') # 生成软标签:0→[0,0.2],1→[0.8,1.0] train_labels = np.where(train_labels == 0, np.random.uniform(0, 0.2, size=train_labels.shape), np.random.uniform(0.8, 1.0, size=train_labels.shape)) test_labels = np.where(test_labels == 0, np.random.uniform(0, 0.2, size=test_labels.shape), np.random.uniform(0.8, 1.0, size=test_labels.shape)) # 构建模型:两种模式二选一 # ----------------模式1:使用sigmoid输出(和原代码一致,兼容软标签)---------------- vgg = vgg16.VGG16(include_top=False, weights='imagenet', input_shape=(32, 32, 3)) output = vgg.layers[-1].output output = layers.Flatten()(output) output = layers.Dense(512, activation='relu')(output) output = layers.Dropout(0.2)(output) output = layers.Dense(256, activation='relu')(output) output = layers.Dropout(0.2)(output) predictions = layers.Dense(units=1, activation="sigmoid")(output) model = Model(inputs=vgg.input, outputs=predictions) # 使用官方BinaryCrossentropy,from_logits=False(默认),兼容软标签 model.compile(optimizer=Adam(learning_rate=.0001), loss=BinaryCrossentropy(), metrics=[soft_label_accuracy]) # ----------------模式2:使用logits输出(数值稳定性更好)---------------- # vgg = vgg16.VGG16(include_top=False, weights='imagenet', input_shape=(32, 32, 3)) # output = vgg.layers[-1].output # output = layers.Flatten()(output) # output = layers.Dense(512, activation='relu')(output) # output = layers.Dropout(0.2)(output) # output = layers.Dense(256, activation='relu')(output) # output = layers.Dropout(0.2)(output) # predictions = layers.Dense(units=1, activation=None)(output) # 去掉sigmoid,输出logits # model = Model(inputs=vgg.input, outputs=predictions) # # 使用from_logits=True,损失函数内部自动做sigmoid转换 # model.compile(optimizer=Adam(learning_rate=.0001), # loss=BinaryCrossentropy(from_logits=True), # metrics=[soft_label_accuracy]) history = model.fit(train_images, train_labels, epochs=100, validation_data=(test_images, test_labels)) plt.plot(history.history['soft_label_accuracy'], label='accuracy') plt.plot(history.history['val_soft_label_accuracy'], label='val_accuracy') plt.xlabel('Epoch') plt.ylabel('Accuracy') plt.ylim([0.5, 1]) plt.legend(loc='lower right') test_loss, test_acc = model.evaluate(test_images, test_labels, verbose=2) print(test_acc)
效果说明
修改后,准确率会回到正常水平(接近原有的98%),此时就能正确验证软标签对分类性能的影响。
内容的提问来源于stack exchange,提问作者IKAHN
相关产品推荐
相关产品推荐

