基于ResNet50的二分类目标检测CNN模型训练后异常输出极值边界框问题求助
基于ResNet50的二分类目标检测CNN模型训练后异常输出极值边界框问题求助
各位大佬好!我最近在做一个二分类目标检测任务,用Keras的ResNet50做迁移学习搭建了CNN模型,但遇到了棘手的问题——不管我训练多少次(已经跑过好几轮100+epoch了),模型输出的边界框坐标总是固定的[0,0,0,1],对应到我300x300的输入图就是Xmin、Ymin、Xmax全为0,Ymax为300,完全不符合真实标注。
先跟大家同步下我的数据情况:
- 输入图像形状是
(batch size, 300, 300, 3) - 边界框标签已经做了0-1归一化处理,形状为
(batch size, 4),打印标注数据能看到正常的结果,比如某一行是[0.74666667 0.32333333 0.88333333 0.69 ] - 分类标签是独热编码格式,形状为
(batch size, 2)
一开始我怀疑是DIoU损失函数的问题,换成MSE损失后,模型还是输出同样的极值边界框,看来不是损失函数的锅。下面是我的模型架构和DIoU损失函数的代码,麻烦各位帮忙排查下问题所在!
模型架构代码
from tensorflow.keras import layers from tensorflow.keras import models from keras.applications.resnet50 import ResNet50 res = ResNet50(weights ='imagenet', include_top = False, input_shape =(300, 300, 3)) x = res.output x = layers.MaxPooling2D((2, 2))(x) # Flatten and Fully connected layers x = layers.BatchNormalization()(x) x = layers.Flatten()(x) x = layers.Dropout(0.5)(x) x = layers.Dense(512, activation='relu')(x) x = layers.BatchNormalization()(x) x = layers.Dropout(0.50)(x) x = layers.Dense(512)(x) # Output layers with proper names bbox_output = layers.Dense(4, activation='sigmoid', name='bbox_output')(x) # Bounding box output class_output = layers.Dense(2, activation='softmax', name='class_output')(x) # Class output # Define the model model = models.Model(inputs = res.input, outputs=[bbox_output, class_output]) model.compile( optimizer='Adam', loss={'bbox_output': diou_loss, 'class_output': 'categorical_crossentropy'}, metrics={'bbox_output': 'MSE', 'class_output': 'accuracy'}) callbacks = [ keras.callbacks.EarlyStopping( # Stop training when `val_loss` is no longer improving monitor='val_bbox_output_loss', # "no longer improving" being defined as "no better than 1e-2 less" min_delta=1e-2, # "no longer improving" being further defined as "for at least 2 epochs" patience=20, verbose=1, )] history = model.fit(X, [bbox_labels,class_labels], epochs=100,batch_size=50, verbose=1, validation_data=(X_test,[bbox_labels_test,class_labels_test]), callbacks=callbacks)
DIoU损失函数代码
def diou_loss(y_true, y_pred, epsilon=1e-7): # Use fixed image dimensions (256x256) print('yt',y_true) print('yp',y_pred) x_min_inter = tf.maximum(y_true[..., 0], y_pred[..., 0]) y_min_inter = tf.maximum(y_true[..., 1], y_pred[..., 1]) x_max_inter = tf.minimum(y_true[..., 2], y_pred[..., 2]) y_max_inter = tf.minimum(y_true[..., 3], y_pred[..., 3]) inter_area = tf.maximum(0.0, x_max_inter - x_min_inter) * tf.maximum(0.0, y_max_inter - y_min_inter) print('ia',inter_area) true_area = tf.maximum(0.0, y_true[..., 2] - y_true[..., 0]) * tf.maximum(0.0, y_true[..., 3] - y_true[..., 1]) pred_area = tf.maximum(0.0, y_pred[..., 2] - y_pred[..., 0]) * tf.maximum(0.0, y_pred[..., 3] - y_pred[..., 1]) union_area = true_area + pred_area - inter_area print('ua',union_area) iou = inter_area / tf.maximum(union_area, epsilon) print('iou',iou) # Calculate the center coordinates of the true and predicted boxes true_center_x = (y_true[..., 0] + y_true[..., 2]) / 2.0 true_center_y = (y_true[..., 1] + y_true[..., 3]) / 2.0 pred_center_x = (y_pred[..., 0] + y_pred[..., 2]) / 2.0 pred_center_y = (y_pred[..., 1] + y_pred[..., 3]) / 2.0 # Calculate the squared Euclidean distance between the centers center_distance = (true_center_x - pred_center_x) ** 2 + (true_center_y - pred_center_y) ** 2 print('cd',center_distance) # Calculate the coordinates of the smallest enclosing box x_min_enclosing = tf.minimum(y_true[..., 0], y_pred[..., 0]) y_min_enclosing = tf.minimum(y_true[..., 1], y_pred[..., 1]) x_max_enclosing = tf.maximum(y_true[..., 2], y_pred[..., 2]) y_max_enclosing = tf.maximum(y_true[..., 3], y_pred[..., 3]) # Calculate the diagonal length squared of the enclosing box enclosing_diagonal = (x_max_enclosing - x_min_enclosing) ** 2 + (y_max_enclosing - y_min_enclosing) ** 2 print('ed',enclosing_diagonal) # Calculate the DIoU # Return the DIoU loss return 1.0 - iou + ((center_distance) / tf.maximum(enclosing_diagonal, 1e-7))
我已经卡了好几天了,实在找不到问题出在哪,恳请各位大佬给点思路,谢谢大家!
备注:内容来源于stack exchange,提问作者Caleb Joseph
相关产品推荐
相关产品推荐

