You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用VNCoreMLRequest同时检测多个对象?

如何用VNCoreMLRequest实现多对象同时检测?

看你已经把VNCoreMLRequest的基础框架搭起来了,要实现多对象同时检测其实只要调整几个关键地方就行,我给你一步步拆解:

首先,换个支持多对象检测的模型

你现在用的Resnet50是图像分类模型,只能识别整张图的主类别,没法同时检测多个对象。得换成目标检测类的CoreML模型,比如YOLO系列、MobileNet SSD,或者Apple官方适配的YOLOv8n这类模型。

然后,优化代码结构并正确配置请求

先把模型初始化的逻辑移到帧捕获外面(别每次抓帧都重新创建模型,太耗性能),再补全VNCoreMLRequest的配置,添加结果回调来解析多对象数据:

extension CameraVC : AVCaptureVideoDataOutputSampleBufferDelegate {
    // 懒加载初始化检测模型,只执行一次
    private lazy var detectionModel: VNCoreMLModel? = {
        do {
            // 替换成你的目标检测模型,比如YOLOv8n
            let coreMLModel = try YOLOv8n(configuration: MLModelConfiguration()).model
            return try VNCoreMLModel(for: coreMLModel)
        } catch {
            print("模型初始化失败:\(error.localizedDescription)")
            return nil
        }
    }()

    func captureOutput(_ output: AVCaptureOutput, didOutput sampleBuffer: CMSampleBuffer, from connection: AVCaptureConnection) {
        print("Camera was able to capture a frame: ", Date())
        guard let pixelBuffer = CMSampleBufferGetImageBuffer(sampleBuffer) else { return }
        guard let model = detectionModel else { return }

        // 初始化VNCoreMLRequest,设置结果回调
        let detectionRequest = VNCoreMLRequest(model: model) { [weak self] request, error in
            guard let self = self else { return }
            
            if let error = error {
                print("检测请求出错:\(error.localizedDescription)")
                return
            }

            // 解析多对象检测结果
            guard let results = request.results as? [VNRecognizedObjectObservation] else {
                print("无法解析检测结果")
                return
            }

            // 遍历所有检测到的对象
            for object in results {
                // 过滤低置信度结果(比如只保留置信度>0.5的)
                guard let topLabel = object.labels.first, topLabel.confidence > 0.5 else { continue }
                
                print("检测到:\(topLabel.identifier),置信度:\(String(format: "%.2f", topLabel.confidence))")
                print("位置(归一化坐标):\(object.boundingBox)")

                // 如果要在预览画面画框,必须回到主线程操作
                DispatchQueue.main.async {
                    // 这里可以添加绘制检测框的逻辑
                    self.drawBoundingBox(for: object)
                }
            }
        }

        // 设置图像缩放选项,适配模型输入要求
        detectionRequest.imageCropAndScaleOption = .scaleFill

        // 用全局队列执行检测,避免阻塞相机回调
        let imageHandler = VNImageRequestHandler(cvPixelBuffer: pixelBuffer, options: [:])
        DispatchQueue.global(qos: .userInitiated).async {
            do {
                try imageHandler.perform([detectionRequest])
            } catch {
                print("执行检测请求失败:\(error.localizedDescription)")
            }
        }
    }

    // 可选:添加绘制检测框的方法
    private func drawBoundingBox(for observation: VNRecognizedObjectObservation) {
        let previewSize = previewView.bounds.size
        // 转换归一化坐标到UIKit坐标系(CoreML的原点在左下角,UIKit在左上角)
        let x = observation.boundingBox.origin.x * previewSize.width
        let y = (1 - observation.boundingBox.origin.y - observation.boundingBox.size.height) * previewSize.height
        let width = observation.boundingBox.size.width * previewSize.width
        let height = observation.boundingBox.size.height * previewSize.height

        let boundingBox = CGRect(x: x, y: y, width: width, height: height)

        // 这里可以创建CAShapeLayer来绘制矩形框,添加到previewView的layer上
        let boxLayer = CAShapeLayer()
        boxLayer.path = UIBezierPath(rect: boundingBox).cgPath
        boxLayer.strokeColor = UIColor.red.cgColor
        boxLayer.fillColor = UIColor.clear.cgColor
        boxLayer.lineWidth = 2.0
        previewView.layer.addSublayer(boxLayer)

        // 记得要移除旧的图层,避免叠加太多
        // 可以在每次画框前先清除previewView上的所有boxLayer
    }
}

最后,几个关键注意事项

  • 模型转换:如果你的模型是PyTorch/TensorFlow格式,需要用CoreML Tools转换成CoreML格式,确保输出是VNRecognizedObjectObservation类型。
  • 性能优化:用懒加载初始化模型、把检测任务放到全局队列、设置置信度阈值过滤无效结果,这些都能提升实时检测的流畅度。
  • 坐标系转换:CoreML返回的boundingBox是归一化坐标,且原点在左下角,和UIKit的坐标系不一样,必须转换后才能正确画框。

内容的提问来源于stack exchange,提问作者Barkley

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:24:22