You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++中迁移YoloV4到TensorRT遇检测结果解析难题求助

Yolov4 TensorRT推理结果解析异常问题解决

问题描述

我之前用YoloV4/C++/OpenCV做目标检测,运行正常,现在为了提升性能迁移到NVIDIA TensorRT,已经把.weights转成ONNX再转成TensorRT engine(转换代码如下),但推理后解析检测结果时遇到问题:得到的类别概率极小(<1E-05),bounding box坐标异常(极小或负值)。

我的Yolov4配置:

  • 20个类别
  • 输入尺寸608x608x3
  • 9个anchors:12,16,19,36,40,28,36,75,76,55,72,146,142,110,192,243,459,401

推理输出两个缓冲区:

  • 1x22743x1x4:推测是bounding box坐标
  • 1x22743x20:推测是类别概率

我想知道:22743这个数值怎么来的?怎么正确解析输出得到正常的坐标和类别概率?

转换代码

void ONNXConvert()
{
    MyLogger logger;
    nvinfer1::IBuilder* builder = nvinfer1::createInferBuilder(logger);
    nvinfer1::INetworkDefinition* network = builder->createNetworkV2(1U << static_cast<int>(nvinfer1::NetworkDefinitionCreationFlag::kEXPLICIT_BATCH));

    // Load ONNX model
    const auto parser = nvonnxparser::createParser(*network, logger);

    // Parse the ONNX model
    // Some code here...

    std::ifstream onnxFile(onnxModelFile, std::ios::binary);
    if (!onnxFile)
    {
        std::cerr << "Error opening ONNX model file. " << onnxModelFile << std::endl;
        return;
    }
    onnxFile.seekg(0, onnxFile.end);
    const size_t modelSize = onnxFile.tellg();
    onnxFile.seekg(0, onnxFile.beg);

    // Allocate buffer to hold the ONNX model
    std::vector<char> onnxModelBuffer(modelSize);
    onnxFile.read(onnxModelBuffer.data(), modelSize);

    if (!parser->parse(onnxModelBuffer.data(), modelSize))
    {
        std::cerr << "Error parsing ONNX model." << std::endl;
        return;
    }

    // Create a builder configuration
    nvinfer1::IBuilderConfig* config = builder->createBuilderConfig();

    // Set configuration options as needed
    config->setMemoryPoolLimit(nvinfer1::MemoryPoolType::kWORKSPACE, 1 << 30);

    nvinfer1::IHostMemory* serializedEngine = builder->buildSerializedNetwork(*network, *config);
    std::cout << "Number of layers in the network: " << network->getNbLayers() << std::endl;
    std::ofstream outFile("yolov4.engine", std::ios::binary);
    outFile.write(reinterpret_cast<const char*>(serializedEngine->data()), serializedEngine->size());
    outFile.close();

    builder->destroy();
    network->destroy();
    serializedEngine->destroy();
}

我之前尝试的解析代码

for (int d = 0; d < 22743; d++)
{
    float maxProb = -1000.0f;
    int classId = -1;
    for (int c = 0; c < 20; c++)
    {
        if (classes[d * 20 + c] > maxProb)
        {
            maxProb = classes[d * 20 + c];
            classId = c;
        }
    }

    if (maxProb > CONFIDENCE_THRESHOLD)
    {
        float boxX = boxes[d * 4];
        float boxY = boxes[d * 4 + 1];
        float boxW = boxes[d * 4 + 2];
        float boxH = boxes[d * 4 + 3];
    }
}

解决方法

一、22743的计算逻辑

Yolov4是多尺度检测模型,会在3个不同尺度的特征图上生成检测框:

  • 大尺度特征图:608/32 = 19x19,对应3个大anchors(142,110,192,243,459,401),总框数:19*19*3 = 1083
  • 中尺度特征图:608/16 = 38x38,对应3个中anchors(36,75,76,55,72,146),总框数:38*38*3 = 4332
  • 小尺度特征图:608/8 = 76x76,对应3个小anchors(12,16,19,36,40,28),总框数:76*76*3 = 17328

三者相加:1083 + 4332 + 17328 = 22743,这就是输出的检测框总数。

二、正确的解析步骤

Yolov4的输出是原始预测值,需要经过激活函数和坐标转换才能得到可用结果:

1. 类别概率计算

输出的1x22743x20是未经过sigmoid激活的logits,需要先转换为0-1之间的概率:

float sigmoid(float x) {
    return 1.0f / (1.0f + exp(-x));
}

同时,每个检测框还有一个目标置信度(部分ONNX模型会把这个值和坐标放在同一个输出张量里),需要同样经过sigmoid激活后,与类别概率相乘得到最终的类别置信度。

2. Bounding Box坐标转换

输出的1x22743x1x4是(tx, ty, tw, th)原始预测值,需要结合对应尺度的网格和anchors转换为真实坐标:

  • bx = sigmoid(tx) + grid_x(grid_x是当前框所在网格的x坐标)
  • by = sigmoid(ty) + grid_y(grid_y是当前框所在网格的y坐标)
  • bw = anchor_w * exp(tw)(anchor_w是当前框对应的anchor宽度)
  • bh = anchor_h * exp(th)(anchor_h是当前框对应的anchor高度)

再将这些值映射到输入图像尺寸:

  • 最终x坐标:bx * (608 / 特征图宽度)
  • 最终y坐标:by * (608 / 特征图高度)
  • 最终宽度:bw * (608 / 特征图宽度)
  • 最终高度:bh * (608 / 特征图高度)

如果需要左上角/右下角坐标,可进一步转换:

  • x1 = bx - bw/2
  • y1 = by - bh/2
  • x2 = bx + bw/2
  • y2 = by + bh/2

3. 分尺度匹配anchors

22743个框需要按三个尺度拆分,分别对应各自的anchors组:

  • 前17328个:对应76x76尺度,使用前3组anchors
  • 接下来4332个:对应38x38尺度,使用中间3组anchors
  • 最后1083个:对应19x19尺度,使用最后3组anchors

三、修正后的解析代码示例

#include <cmath>
#include <vector>

const int INPUT_WIDTH = 608;
const int INPUT_HEIGHT = 608;
const int NUM_CLASSES = 20;
const float CONF_THRESH = 0.5f;
const float NMS_THRESH = 0.4f;

// Anchors分组:小、中、大尺度对应
std::vector<std::vector<float>> anchors = {
    {12,16}, {19,36}, {40,28},    // 76x76尺度
    {36,75}, {76,55}, {72,146},   // 38x38尺度
    {142,110}, {192,243}, {459,401} // 19x19尺度
};

float sigmoid(float x) {
    return 1.0f / (1.0f + exp(-x));
}

struct Detection {
    float x1, y1, x2, y2;
    float prob;
    int classId;
};

void processScale(float* boxes, float* classes, int startIdx, int numBoxes, int gridW, int gridH, 
                  std::vector<std::vector<float>>::iterator anchorStart, 
                  std::vector<std::vector<float>>::iterator anchorEnd,
                  std::vector<Detection>& detections) {
    float strideW = INPUT_WIDTH / (float)gridW;
    float strideH = INPUT_HEIGHT / (float)gridH;

    for (int i = 0; i < numBoxes; i++) {
        int idx = startIdx + i;

        // 读取原始预测值
        float tx = boxes[idx*4];
        float ty = boxes[idx*4 + 1];
        float tw = boxes[idx*4 + 2];
        float th = boxes[idx*4 + 3];
        // 注意:如果你的ONNX模型把目标置信度放在坐标输出里,这里需要读取boxes[idx*4 +4]

        // 计算当前框的网格位置和对应anchor
        int gridX = i % gridW;
        int gridY = (i / gridW) % gridH;
        int anchorIdx = (i / (gridW * gridH)) % 3;
        auto& anchor = *(anchorStart + anchorIdx);
        float anchorW = anchor[0];
        float anchorH = anchor[1];

        // 转换为真实坐标
        float bx = (sigmoid(tx) + gridX) * strideW;
        float by = (sigmoid(ty) + gridY) * strideH;
        float bw = anchorW * exp(tw);
        float bh = anchorH * exp(th);

        // 转换为左上角/右下角坐标
        float x1 = bx - bw/2;
        float y1 = by - bh/2;
        float x2 = bx + bw/2;
        float y2 = by + bh/2;

        // 计算目标置信度(如果模型输出包含该值)
        float objConf = sigmoid(boxes[idx*4 +4]); // 根据你的模型结构调整

        // 计算最大类别概率
        float maxProb = 0.0f;
        int classId = -1;
        for (int c = 0; c < NUM_CLASSES; c++) {
            float prob = sigmoid(classes[idx*NUM_CLASSES + c]);
            if (prob > maxProb) {
                maxProb = prob;
                classId = c;
            }
        }
        float finalConf = objConf * maxProb;

        // 过滤低置信度框
        if (finalConf > CONF_THRESH) {
            Detection det;
            det.x1 = x1;
            det.y1 = y1;
            det.x2 = x2;
            det.y2 = y2;
            det.prob = finalConf;
            det.classId = classId;
            detections.push_back(det);
        }
    }
}

std::vector<Detection> parseOutputs(float* boxes, float* classes) {
    std::vector<Detection> detections;

    // 处理三个尺度的检测框
    processScale(boxes, classes, 0, 17328, 76, 76, anchors.begin(), anchors.begin()+3, detections);
    processScale(boxes, classes, 17328, 4332, 38, 38, anchors.begin()+3, anchors.begin()+6, detections);
    processScale(boxes, classes, 17328+4332, 1083, 19, 19, anchors.begin()+6, anchors.end(), detections);

    // 可选:添加非极大值抑制去除重复框
    // applyNMS(detections, NMS_THRESH);

    return detections;
}

注意事项

  • 确认ONNX模型输出结构:部分Yolov4 ONNX模型会把坐标+目标置信度+类别logits放在同一个输出张量(形状1x22743x(4+1+20)),需根据实际结构调整代码。
  • 输入预处理一致性:确保推理时的输入归一化、通道顺序(RGB/BGR)和训练时完全一致,否则会导致结果异常。
  • 非极大值抑制(NMS):解析完成后,建议添加NMS步骤去除重叠度高的重复检测框。

内容的提问来源于stack exchange,提问作者Stef

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 02:48:12