求助:Mediapipe手掌检测模型输出转Bounding Box的计算方法
手掌检测模型输出转Bounding Box的计算步骤
Mediapipe手掌检测模型基于锚点(Anchor)机制,输出的dx, dy, w, h是相对于预设锚点的偏移量,需结合锚点参数、激活函数转换及坐标映射得到真实Bounding Box,具体步骤如下:
1. 筛选有效锚点
- 模型输出的置信度是未激活的原始值(如示例中的
-2090.7869),先通过Sigmoid函数转换为0-1区间的置信度:float confidence = 1.0f / (1.0f + exp(-raw_confidence)); - 设置置信度阈值(建议0.7),仅保留置信度高于阈值的锚点,减少无效计算。
2. 获取预设锚点参数
模型的2016个锚点是提前按多尺度特征图生成的,每个锚点包含:
anchor_x, anchor_y:锚点中心在特征图上的坐标anchor_w, anchor_h:锚点的预设宽高
锚点参数可从Mediapipe手掌检测的管线配置或源码中提取,特征图尺寸与模型输入图像尺寸的比例固定(比如输入192x192时,特征图由多个尺度层组成,总锚点数为2016)。
3. 转换偏移量为特征图上的bbox
模型输出的dx, dy, w, h需通过以下公式转换:
- 对
dx, dy应用Sigmoid,得到相对于锚点中心的偏移比例:float dx_sigmoid = 1.0f / (1.0f + exp(-dx)); float dy_sigmoid = 1.0f / (1.0f + exp(-dy)); - 计算特征图上的bbox中心坐标:
float box_center_x = anchor_x + dx_sigmoid * anchor_w; float box_center_y = anchor_y + dy_sigmoid * anchor_h; - 对
w, h应用指数函数,得到真实宽高:float box_w = anchor_w * exp(w); float box_h = anchor_h * exp(h);
4. 映射到原始图像尺寸
假设模型输入图像尺寸为input_w × input_h,原始图像尺寸为orig_w × orig_h:
- 先将特征图上的bbox转换为输入图像坐标:
计算特征图与输入图像的缩放比例scale_w = input_w / feature_map_total_width,scale_h = input_h / feature_map_total_height(该比例可从锚点生成逻辑中推导)float input_center_x = box_center_x * scale_w; float input_center_y = box_center_y * scale_h; float input_box_w = box_w * scale_w; float input_box_h = box_h * scale_h; - 再将输入图像的bbox映射到原始图像(若输入图像是从原始图像缩放得到):
float scale_orig_w = (float)orig_w / input_w; float scale_orig_h = (float)orig_h / input_h; float orig_center_x = input_center_x * scale_orig_w; float orig_center_y = input_center_y * scale_orig_h; float orig_box_w = input_box_w * scale_orig_w; float orig_box_h = input_box_h * scale_orig_h;
5. 计算bbox的四角坐标
通过中心坐标和宽高推导左上角、右下角坐标:
float x1 = orig_center_x - orig_box_w / 2.0f; float y1 = orig_center_y - orig_box_h / 2.0f; float x2 = orig_center_x + orig_box_w / 2.0f; float y2 = orig_center_y + orig_box_h / 2.0f;
6. 非极大值抑制(NMS)
多个锚点可能检测到同一手掌,需用NMS去除重叠bbox:
- 计算有效bbox之间的IOU(交并比)
- 保留IOU低于阈值(建议0.3)的最高置信度bbox
内容的提问来源于stack exchange,提问作者M.Akyuzlu
相关产品推荐
相关产品推荐

