基于VGG16的目标检测:为Bounding Box添加概率值遇问题
为双目标Bounding Box添加置信度的解决方案
首先明确:decode_predictions()是专门为ImageNet预训练分类模型设计的,要求输入是1000维的分类概率数组,你的模型输出是Bounding Box坐标,完全不匹配,所以报错是必然的——这个函数对你的场景毫无用处。
下面给出无需额外训练模型的解决方案,直接在现有架构中扩展置信度输出:
方案1:拆分坐标与置信度为独立输出分支
1. 修改输出层构建函数
让每个目标分支同时输出Bounding Box坐标和置信度(置信度代表目标存在的概率,用sigmoid归一化到0-1区间):
def output_bbox_with_confidence(name_prefix): flatten_out = Flatten()(vgg.output) # 共享全连接层提取特征 shared_dense = Dense(512, activation="relu")(flatten_out) shared_dense = Dense(256, activation="relu")(shared_dense) shared_dense = Dense(128, activation="relu")(shared_dense) shared_dense = Dense(64, activation="relu")(shared_dense) shared_dense = Dense(32, activation="relu")(shared_dense) # Bounding Box坐标输出(4个值,sigmoid归一化到0-1) bbox = Dense(4, activation="sigmoid", name=f"{name_prefix}_bbox")(shared_dense) # 置信度输出(1个值,sigmoid归一化到0-1) confidence = Dense(1, activation="sigmoid", name=f"{name_prefix}_conf")(shared_dense) return bbox, confidence
2. 构建多输出模型
调用修改后的函数生成两个目标的输出分支:
# 生成第一个目标的bbox和置信度 bbox1, conf1 = output_bbox_with_confidence("target1") # 生成第二个目标的bbox和置信度 bbox2, conf2 = output_bbox_with_confidence("target2") # 定义模型输入输出 model = Model(inputs=vgg.input, outputs=[bbox1, conf1, bbox2, conf2])
3. 调整损失函数与编译
为不同输出分支匹配对应损失:
- 坐标分支:用均方误差(MSE)或平滑L1损失(适合回归任务)
- 置信度分支:用二元交叉熵(二分类:目标存在/不存在)
示例编译代码:
from tensorflow.keras.losses import MeanSquaredError, BinaryCrossentropy losses = { "target1_bbox": MeanSquaredError(), "target1_conf": BinaryCrossentropy(), "target2_bbox": MeanSquaredError(), "target2_conf": BinaryCrossentropy() } # 给不同损失分配权重,可根据实际效果调整 lossWeights = { "target1_bbox": 1.0, "target1_conf": 0.5, "target2_bbox": 1.0, "target2_conf": 0.5 } model.compile(optimizer=Adam(), loss=losses, loss_weights=lossWeights)
4. 调整训练数据格式
trainTargets和testTargets需要对应模型的4个输出,可使用字典或列表格式:
- 字典格式:
{"target1_bbox": train_bbox1, "target1_conf": train_conf1, "target2_bbox": train_bbox2, "target2_conf": train_conf2} - 列表格式:
[train_bbox1, train_conf1, train_bbox2, train_conf2]
其中train_conf1是每个样本对应第一个目标是否存在的标签(0或1,或标注的置信度),train_bbox1是对应的4个坐标值。
5. 预测时获取置信度
训练完成后,直接从预测结果中提取置信度:
# 单帧预测 predictions = model.predict(frame_input) # 结果顺序对应模型输出:bbox1, conf1, bbox2, conf2 bbox1 = predictions[0][0] conf1 = predictions[1][0][0] # 提取第一个目标的置信度 bbox2 = predictions[2][0] conf2 = predictions[3][0][0] # 提取第二个目标的置信度 print(f"目标1 BBox: {bbox1}, 置信度: {conf1:.4f}") print(f"目标2 BBox: {bbox2}, 置信度: {conf2:.4f}")
方案2:单个分支输出5维向量(坐标+置信度)
如果不想拆分多个输出,可让每个目标分支直接输出5个值(前4个为坐标,最后1个为置信度):
修改输出层函数
def output_bbox_with_confidence(name): flatten_out = Flatten()(vgg.output) bbox = Dense(512,activation="relu")(flatten_out) bbox = Dense(256,activation="relu")(bbox) bbox = Dense(128,activation="relu")(bbox) bbox = Dense(64,activation="relu")(bbox) bbox = Dense(32,activation="relu")(bbox) # 输出5个值:前4个是bbox坐标,最后1个是置信度 bbox = Dense(5, activation="sigmoid", name=name)(bbox) return bbox
自定义损失函数
需要拆分坐标和置信度的损失计算:
from tensorflow.keras.losses import MeanSquaredError, BinaryCrossentropy def custom_bbox_loss(y_true, y_pred): # y_true前4维是坐标,第5维是置信度标签 coord_loss = MeanSquaredError()(y_true[:, :4], y_pred[:, :4]) conf_loss = BinaryCrossentropy()(y_true[:, 4:], y_pred[:, 4:]) return coord_loss + 0.5 * conf_loss # 构建模型 bboxlayer = output_bbox_with_confidence("target1") bboxlayer_2 = output_bbox_with_confidence("target2") model = Model(inputs=vgg.input, outputs=(bboxlayer, bboxlayer_2)) # 编译模型 losses = {"target1": custom_bbox_loss, "target2": custom_bbox_loss} model.compile(optimizer=Adam(), loss=losses, loss_weights=lossWeights)
预测时提取置信度
predictions = model.predict(frame_input) # 第一个目标的结果 target1_pred = predictions[0][0] bbox1 = target1_pred[:4] conf1 = target1_pred[4] # 第二个目标的结果 target2_pred = predictions[1][0] bbox2 = target2_pred[:4] conf2 = target2_pred[4]
内容的提问来源于stack exchange,提问作者Luca
相关产品推荐
相关产品推荐

