Faster R-CNN ResNet50模型维度不匹配及自定义Loss报错求助
解决Faster R-CNN风格口罩检测模型的训练报错问题
问题1:ValueError - 预测与标注维度不匹配
报错原因
模型输出是两个张量:[bbox_output(shape=(?,115,4)), class_output(shape=(?,115,1))],但训练时传入的标注是单一的(?,115,5)张量。使用loss=[bbox_loss, class_loss]时,Keras会将整个标注张量分别传给两个loss函数,导致bbox_loss中y_true(shape=(?,115,5))与y_pred(shape=(?,115,4))维度不匹配。
解决方法
修改数据集处理流程,将标注拆分为边界框张量和类别标签张量两个独立部分,让模型的两个输出分别对应这两个输入:
# 拆分填充后的标注为边界框和标签 def split_annotations(padded_annotations): bboxes = padded_annotations[:, :, :4] labels = padded_annotations[:, :, 4].astype(np.int32) # 转为整数类型适配稀疏交叉熵 return bboxes, labels # 拆分训练/验证标注 train_bboxes, train_labels = split_annotations(padded_train_annotations) val_bboxes, val_labels = split_annotations(padded_val_annotations) # 修改数据集创建函数,返回(图像, (边界框, 标签))格式 def create_tf_dataset(images, bboxes, labels, batch_size): dataset = tf.data.Dataset.from_tensor_slices((images, (bboxes, labels))) dataset = dataset.shuffle(buffer_size=len(images)) dataset = dataset.batch(batch_size) return dataset # 生成新的训练/验证数据集 train_dataset = create_tf_dataset(train_images, train_bboxes, train_labels, batch_size) val_dataset = create_tf_dataset(val_images, val_bboxes, val_labels, batch_size) # 验证数据集形状 for img_batch, (bbox_batch, label_batch) in train_dataset.take(1): print("Image batch shape:", img_batch.shape) print("Bbox batch shape:", bbox_batch.shape) print("Label batch shape:", label_batch.shape)
同时修正loss函数,加入有效框掩码(跳过填充的全0框):
def bbox_loss(y_true, y_pred): # 只计算非填充框的损失 mask = tf.reduce_any(y_true != 0, axis=-1) mask = tf.cast(mask, tf.float32) loss = tf.losses.mean_squared_error(y_true, y_pred) return tf.reduce_sum(loss * mask) / tf.maximum(tf.reduce_sum(mask), 1.0) def class_loss(y_true, y_pred): # 使用稀疏交叉熵处理整数标签 return tf.reduce_mean(tf.losses.sparse_categorical_crossentropy(y_true, y_pred))
最后编译模型:
model.compile( optimizer=tf.keras.optimizers.Adam(learning_rate=0.001), loss=[bbox_loss, class_loss], metrics={'bbox_output': 'mse', 'class_output': 'accuracy'} )
问题2:OperatorNotAllowedInGraphError - 无法迭代符号张量
报错原因
在自定义loss函数中直接拆分y_pred为pred_bboxes, pred_labels = y_pred,但TensorFlow图模式下,符号张量不能像Python列表一样迭代拆分,AutoGraph无法处理该操作,因此报错。
解决方法
优先采用上述拆分标注+多输出loss的方案。如果一定要使用联合loss,需修改模型输出为合并张量,并用tf.split在loss函数中拆分:
# 修改模型输出为合并张量 def create_model(backbone): inputs = keras.Input(shape=(512, 512, 3)) x = backbone(inputs, training=False) x = layers.GlobalAveragePooling2D()(x) x = layers.Dense(512, activation='relu')(x) # 边界框输出:115个框,每个框4个坐标 bbox_output = layers.Dense(115 * 4, activation='sigmoid')(x) bbox_output = layers.Reshape((115, 4))(bbox_output) # 分类输出:115个框,每个框3类概率(对应3种口罩状态) class_output = layers.Dense(115 * 3, activation='softmax')(x) class_output = layers.Reshape((115, 3))(class_output) # 合并输出张量 combined_output = layers.concatenate([bbox_output, class_output], axis=-1) model = keras.Model(inputs=inputs, outputs=combined_output) return model # 联合loss函数,用tf.split拆分张量 def custom_loss(y_true, y_pred): # 拆分预测结果为边界框和分类概率 pred_bboxes, pred_labels = tf.split(y_pred, [4, 3], axis=-1) # 拆分真实标注 true_bboxes = y_true[:, :, :4] true_labels = y_true[:, :, 4] # 计算边界框损失(带掩码) mask = tf.reduce_any(true_bboxes != 0, axis=-1) mask = tf.cast(mask, tf.float32) bbox_loss = tf.reduce_sum(tf.losses.mean_squared_error(true_bboxes, pred_bboxes) * mask) / tf.maximum(tf.reduce_sum(mask), 1.0) # 计算分类损失(只计算有效框) class_loss = tf.reduce_sum(tf.losses.sparse_categorical_crossentropy(true_labels, pred_labels) * mask) / tf.maximum(tf.reduce_sum(mask), 1.0) return bbox_loss + class_loss # 编译模型 model.compile( optimizer=tf.keras.optimizers.Adam(learning_rate=0.001), loss=custom_loss )
额外优化建议
- 有效框掩码:训练时跳过填充的全0框,避免引入噪声,上述loss函数已加入该处理。
- 模型结构修正:当前结构并非标准Faster R-CNN,缺少区域建议网络(RPN)、ROI池化等核心组件。若要实现标准Faster R-CNN,建议使用TensorFlow Object Detection API或官方实现示例。
- 骨干网络微调:训练后期可解冻ResNet50的部分层,提升模型性能。
内容的提问来源于stack exchange,提问作者Al-Amin Baba
相关产品推荐
相关产品推荐

