C++中迁移YoloV4到TensorRT遇检测结果解析难题求助
Yolov4 TensorRT推理结果解析异常问题解决
问题描述
我之前用YoloV4/C++/OpenCV做目标检测,运行正常,现在为了提升性能迁移到NVIDIA TensorRT,已经把.weights转成ONNX再转成TensorRT engine(转换代码如下),但推理后解析检测结果时遇到问题:得到的类别概率极小(<1E-05),bounding box坐标异常(极小或负值)。
我的Yolov4配置:
- 20个类别
- 输入尺寸608x608x3
- 9个anchors:
12,16,19,36,40,28,36,75,76,55,72,146,142,110,192,243,459,401
推理输出两个缓冲区:
1x22743x1x4:推测是bounding box坐标1x22743x20:推测是类别概率
我想知道:22743这个数值怎么来的?怎么正确解析输出得到正常的坐标和类别概率?
转换代码
void ONNXConvert() { MyLogger logger; nvinfer1::IBuilder* builder = nvinfer1::createInferBuilder(logger); nvinfer1::INetworkDefinition* network = builder->createNetworkV2(1U << static_cast<int>(nvinfer1::NetworkDefinitionCreationFlag::kEXPLICIT_BATCH)); // Load ONNX model const auto parser = nvonnxparser::createParser(*network, logger); // Parse the ONNX model // Some code here... std::ifstream onnxFile(onnxModelFile, std::ios::binary); if (!onnxFile) { std::cerr << "Error opening ONNX model file. " << onnxModelFile << std::endl; return; } onnxFile.seekg(0, onnxFile.end); const size_t modelSize = onnxFile.tellg(); onnxFile.seekg(0, onnxFile.beg); // Allocate buffer to hold the ONNX model std::vector<char> onnxModelBuffer(modelSize); onnxFile.read(onnxModelBuffer.data(), modelSize); if (!parser->parse(onnxModelBuffer.data(), modelSize)) { std::cerr << "Error parsing ONNX model." << std::endl; return; } // Create a builder configuration nvinfer1::IBuilderConfig* config = builder->createBuilderConfig(); // Set configuration options as needed config->setMemoryPoolLimit(nvinfer1::MemoryPoolType::kWORKSPACE, 1 << 30); nvinfer1::IHostMemory* serializedEngine = builder->buildSerializedNetwork(*network, *config); std::cout << "Number of layers in the network: " << network->getNbLayers() << std::endl; std::ofstream outFile("yolov4.engine", std::ios::binary); outFile.write(reinterpret_cast<const char*>(serializedEngine->data()), serializedEngine->size()); outFile.close(); builder->destroy(); network->destroy(); serializedEngine->destroy(); }
我之前尝试的解析代码
for (int d = 0; d < 22743; d++) { float maxProb = -1000.0f; int classId = -1; for (int c = 0; c < 20; c++) { if (classes[d * 20 + c] > maxProb) { maxProb = classes[d * 20 + c]; classId = c; } } if (maxProb > CONFIDENCE_THRESHOLD) { float boxX = boxes[d * 4]; float boxY = boxes[d * 4 + 1]; float boxW = boxes[d * 4 + 2]; float boxH = boxes[d * 4 + 3]; } }
解决方法
一、22743的计算逻辑
Yolov4是多尺度检测模型,会在3个不同尺度的特征图上生成检测框:
- 大尺度特征图:608/32 = 19x19,对应3个大anchors(
142,110,192,243,459,401),总框数:19*19*3 = 1083 - 中尺度特征图:608/16 = 38x38,对应3个中anchors(
36,75,76,55,72,146),总框数:38*38*3 = 4332 - 小尺度特征图:608/8 = 76x76,对应3个小anchors(
12,16,19,36,40,28),总框数:76*76*3 = 17328
三者相加:1083 + 4332 + 17328 = 22743,这就是输出的检测框总数。
二、正确的解析步骤
Yolov4的输出是原始预测值,需要经过激活函数和坐标转换才能得到可用结果:
1. 类别概率计算
输出的1x22743x20是未经过sigmoid激活的logits,需要先转换为0-1之间的概率:
float sigmoid(float x) { return 1.0f / (1.0f + exp(-x)); }
同时,每个检测框还有一个目标置信度(部分ONNX模型会把这个值和坐标放在同一个输出张量里),需要同样经过sigmoid激活后,与类别概率相乘得到最终的类别置信度。
2. Bounding Box坐标转换
输出的1x22743x1x4是(tx, ty, tw, th)原始预测值,需要结合对应尺度的网格和anchors转换为真实坐标:
bx = sigmoid(tx) + grid_x(grid_x是当前框所在网格的x坐标)by = sigmoid(ty) + grid_y(grid_y是当前框所在网格的y坐标)bw = anchor_w * exp(tw)(anchor_w是当前框对应的anchor宽度)bh = anchor_h * exp(th)(anchor_h是当前框对应的anchor高度)
再将这些值映射到输入图像尺寸:
- 最终x坐标:
bx * (608 / 特征图宽度) - 最终y坐标:
by * (608 / 特征图高度) - 最终宽度:
bw * (608 / 特征图宽度) - 最终高度:
bh * (608 / 特征图高度)
如果需要左上角/右下角坐标,可进一步转换:
x1 = bx - bw/2y1 = by - bh/2x2 = bx + bw/2y2 = by + bh/2
3. 分尺度匹配anchors
22743个框需要按三个尺度拆分,分别对应各自的anchors组:
- 前17328个:对应76x76尺度,使用前3组anchors
- 接下来4332个:对应38x38尺度,使用中间3组anchors
- 最后1083个:对应19x19尺度,使用最后3组anchors
三、修正后的解析代码示例
#include <cmath> #include <vector> const int INPUT_WIDTH = 608; const int INPUT_HEIGHT = 608; const int NUM_CLASSES = 20; const float CONF_THRESH = 0.5f; const float NMS_THRESH = 0.4f; // Anchors分组:小、中、大尺度对应 std::vector<std::vector<float>> anchors = { {12,16}, {19,36}, {40,28}, // 76x76尺度 {36,75}, {76,55}, {72,146}, // 38x38尺度 {142,110}, {192,243}, {459,401} // 19x19尺度 }; float sigmoid(float x) { return 1.0f / (1.0f + exp(-x)); } struct Detection { float x1, y1, x2, y2; float prob; int classId; }; void processScale(float* boxes, float* classes, int startIdx, int numBoxes, int gridW, int gridH, std::vector<std::vector<float>>::iterator anchorStart, std::vector<std::vector<float>>::iterator anchorEnd, std::vector<Detection>& detections) { float strideW = INPUT_WIDTH / (float)gridW; float strideH = INPUT_HEIGHT / (float)gridH; for (int i = 0; i < numBoxes; i++) { int idx = startIdx + i; // 读取原始预测值 float tx = boxes[idx*4]; float ty = boxes[idx*4 + 1]; float tw = boxes[idx*4 + 2]; float th = boxes[idx*4 + 3]; // 注意:如果你的ONNX模型把目标置信度放在坐标输出里,这里需要读取boxes[idx*4 +4] // 计算当前框的网格位置和对应anchor int gridX = i % gridW; int gridY = (i / gridW) % gridH; int anchorIdx = (i / (gridW * gridH)) % 3; auto& anchor = *(anchorStart + anchorIdx); float anchorW = anchor[0]; float anchorH = anchor[1]; // 转换为真实坐标 float bx = (sigmoid(tx) + gridX) * strideW; float by = (sigmoid(ty) + gridY) * strideH; float bw = anchorW * exp(tw); float bh = anchorH * exp(th); // 转换为左上角/右下角坐标 float x1 = bx - bw/2; float y1 = by - bh/2; float x2 = bx + bw/2; float y2 = by + bh/2; // 计算目标置信度(如果模型输出包含该值) float objConf = sigmoid(boxes[idx*4 +4]); // 根据你的模型结构调整 // 计算最大类别概率 float maxProb = 0.0f; int classId = -1; for (int c = 0; c < NUM_CLASSES; c++) { float prob = sigmoid(classes[idx*NUM_CLASSES + c]); if (prob > maxProb) { maxProb = prob; classId = c; } } float finalConf = objConf * maxProb; // 过滤低置信度框 if (finalConf > CONF_THRESH) { Detection det; det.x1 = x1; det.y1 = y1; det.x2 = x2; det.y2 = y2; det.prob = finalConf; det.classId = classId; detections.push_back(det); } } } std::vector<Detection> parseOutputs(float* boxes, float* classes) { std::vector<Detection> detections; // 处理三个尺度的检测框 processScale(boxes, classes, 0, 17328, 76, 76, anchors.begin(), anchors.begin()+3, detections); processScale(boxes, classes, 17328, 4332, 38, 38, anchors.begin()+3, anchors.begin()+6, detections); processScale(boxes, classes, 17328+4332, 1083, 19, 19, anchors.begin()+6, anchors.end(), detections); // 可选:添加非极大值抑制去除重复框 // applyNMS(detections, NMS_THRESH); return detections; }
注意事项
- 确认ONNX模型输出结构:部分Yolov4 ONNX模型会把
坐标+目标置信度+类别logits放在同一个输出张量(形状1x22743x(4+1+20)),需根据实际结构调整代码。 - 输入预处理一致性:确保推理时的输入归一化、通道顺序(RGB/BGR)和训练时完全一致,否则会导致结果异常。
- 非极大值抑制(NMS):解析完成后,建议添加NMS步骤去除重叠度高的重复检测框。
内容的提问来源于stack exchange,提问作者Stef
相关产品推荐
相关产品推荐

