You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Jetson Nano(Jetpack4.5)安装C++版torchvision遇阻,求nms实现

解决Jetson Nano Jetpack 4.5上C++版torchvision::ops::nms的替代实现方案

一、基于LibTorch的手动NMS实现

以下是完全基于LibTorch API实现的NMS函数,无需依赖torchvision库,可直接嵌入你的项目:

#include <torch/torch.h>

torch::Tensor nms(torch::Tensor boxes, torch::Tensor scores, float iou_threshold) {
    // 输入合法性校验
    TORCH_CHECK(boxes.dim() == 2 && boxes.size(1) == 4, "boxes必须是[N, 4]形状的2D张量");
    TORCH_CHECK(scores.dim() == 1 && scores.size(0) == boxes.size(0), "scores必须是与boxes数量匹配的1D张量");
    TORCH_CHECK(iou_threshold >= 0 && iou_threshold <= 1, "iou_threshold取值范围必须在0-1之间");

    if (boxes.numel() == 0) {
        return torch::empty({0}, torch::kLong);
    }

    // 提取框的坐标信息
    auto x1 = boxes.select(1, 0);
    auto y1 = boxes.select(1, 1);
    auto x2 = boxes.select(1, 2);
    auto y2 = boxes.select(1, 3);

    // 计算每个框的面积
    auto area = (x2 - x1 + 1) * (y2 - y1 + 1);
    // 按置信度降序排序,获取索引
    auto sorted_indices = scores.argsort(0, true);

    torch::Tensor keep = torch::empty_like(sorted_indices);
    int64_t keep_size = 0;
    auto sorted_indices_cpu = sorted_indices.cpu();
    auto x1_cpu = x1.cpu();
    auto y1_cpu = y1.cpu();
    auto x2_cpu = x2.cpu();
    auto y2_cpu = y2.cpu();
    auto area_cpu = area.cpu();

    while (sorted_indices_cpu.numel() > 0) {
        // 保留当前置信度最高的框
        int64_t current_idx = sorted_indices_cpu[0].item<int64_t>();
        keep[keep_size++] = current_idx;

        if (sorted_indices_cpu.numel() == 1) {
            break;
        }

        // 计算当前框与剩余框的IOU
        auto rest_indices = sorted_indices_cpu.slice(0, 1);
        auto xx1 = torch::max(x1_cpu[current_idx], x1_cpu.index({rest_indices}));
        auto yy1 = torch::max(y1_cpu[current_idx], y1_cpu.index({rest_indices}));
        auto xx2 = torch::min(x2_cpu[current_idx], x2_cpu.index({rest_indices}));
        auto yy2 = torch::min(y2_cpu[current_idx], y2_cpu.index({rest_indices}));

        auto w = torch::max(torch::tensor(0.0f), xx2 - xx1 + 1);
        auto h = torch::max(torch::tensor(0.0f), yy2 - yy1 + 1);
        auto inter = w * h;
        auto iou = inter / (area_cpu[current_idx] + area_cpu.index({rest_indices}) - inter);

        // 过滤IOU超过阈值的框
        auto mask = iou < iou_threshold;
        sorted_indices_cpu = rest_indices.index({mask});
    }

    return keep.slice(0, 0, keep_size).to(boxes.device());
}

使用说明

  1. 将上述代码加入项目源码,确保已正确链接LibTorch库
  2. 编译时需统一**_GLIBCXX_USE_CXX11_ABI**宏定义:
    • 若你的LibTorch采用CXX11 ABI编译,编译项目时添加-D_GLIBCXX_USE_CXX11_ABI=1
    • 否则添加-D_GLIBCXX_USE_CXX11_ABI=0(Jetpack 4.5默认GCC环境通常使用旧ABI,需与LibTorch版本匹配)

二、直接集成torchvision官方NMS源码方案

若坚持使用官方实现,可从torchvision源码中提取NMS模块手动集成:

  1. 从torchvision代码仓库中获取以下文件:
    • torchvision/csrc/ops/nms.cpp
    • torchvision/csrc/ops/nms.h
  2. 将文件复制到你的项目目录
  3. 编译时配置:
    • 包含LibTorch头文件路径
    • 链接LibTorch核心库
  4. 统一ABI宏定义,解决编译时的语法错误

注意事项

  • Jetpack 4.5系统环境较旧,预编译torchvision whl包兼容性差,手动集成源码是更可靠的方式
  • 若遇编译错误,检查GCC版本(Jetpack 4.5默认GCC 7.5)是否与LibTorch编译版本兼容

内容的提问来源于stack exchange,提问作者Davtag

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 20:20:48