FiftyOne与YOLOX的COCO评估指标为何不匹配?如何对齐?
问题:YOLOX与FiftyOne的COCO评估指标不匹配,如何对齐?
背景
我有一个用于测试目标检测模型的数据集test_dataset,分别使用YOLOX的eval.py脚本和FiftyOne进行COCO风格评估,但两者输出的指标存在明显差异,希望让结果对齐。
YOLOX eval.py评估结果
Average forward time: 23.05 ms, Average NMS time: 2.60 ms, Average inference time: 25.65 ms Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.724 Average Precision (AP) @[ IoU=0.50 | area= all | maxDets=100 ] = 0.957 Average Precision (AP) @[ IoU=0.75 | area= all | maxDets=100 ] = 0.831 Average Precision (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.278 Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.591 Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.810 Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets= 1 ] = 0.535 Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets= 10 ] = 0.755 Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.759 Average Recall (AR) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.349 Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.649 Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.839 per class AP: | class | AP | class | AP | class | AP | |:-------------|:-------|:--------|:-------|:-------------|:-------| | cargo | 59.491 | ferry | 87.701 | fishing boat | 67.328 | | sailing boat | 75.134 | | | | | per class AR: | class | AR | class | AR | class | AR | |:-------------|:-------|:--------|:-------|:-------------|:-------| | cargo | 64.802 | ferry | 89.717 | fishing boat | 71.506 | | sailing boat | 77.748 |
FiftyOne评估结果
precision recall f1-score support cargo 0.76 0.91 0.83 606 ferry 0.97 1.00 0.99 990 fishing boat 0.85 0.96 0.91 332 sailing boat 0.87 0.97 0.92 706 micro avg 0.88 0.97 0.92 2634 macro avg 0.87 0.96 0.91 2634 weighted avg 0.88 0.97 0.92 2634
我使用的FiftyOne脚本
results = dataset.evaluate_detections( "predictions", gt_field="detections", compute_mAP=True, method="coco" ) results.print_report()
指标对齐方案
- 明确指标定义差异:FiftyOne的
print_report()默认输出的是IoU=0.5单阈值下的precision/recall,这和YOLOX的AP不是同一概念——AP是precision-recall曲线下的面积,而单阈值precision/recall是曲线上的一个点。如果要对比YOLOX的AP@0.5:0.95,需要查看FiftyOne结果中的results.mAP值,而非print_report输出内容。 - 匹配核心评估参数:
- 确保两者的NMS参数一致:YOLOX eval.py默认NMS IoU阈值为0.65,maxDets为100;若FiftyOne导入的预测结果未经过NMS,需先调用
dataset["predictions"] = dataset["predictions"].nms(iou_thresh=0.65, max_dets=100)处理。 - 对齐IoU阈值范围:COCO标准mAP使用10个IoU阈值(0.5到0.95,步长0.05),FiftyOne需显式指定该范围以匹配YOLOX,修改脚本为:
results = dataset.evaluate_detections( "predictions", gt_field="detections", compute_mAP=True, method="coco", iou_threshes=[0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95], max_dets=100 )
- 确保两者的NMS参数一致:YOLOX eval.py默认NMS IoU阈值为0.65,maxDets为100;若FiftyOne导入的预测结果未经过NMS,需先调用
- 校验数据格式与类别映射:
- 确认YOLOX和FiftyOne使用的bbox坐标格式一致(均为xyxy或xywh),YOLOX默认输出xyxy格式,FiftyOne需对应。
- 确保两者的类别名称、ID完全匹配,无遗漏或错位。
- 统一评估流程:YOLOX eval.py会自动过滤低置信度预测(默认置信度阈值0.001),需确保FiftyOne中的预测结果也过滤掉相同置信度阈值的框,避免无效框干扰评估。
内容的提问来源于stack exchange,提问作者Benoît Boidin
相关产品推荐
相关产品推荐

