如何使用VNCoreMLRequest同时检测多个对象?
如何用VNCoreMLRequest实现多对象同时检测?
看你已经把VNCoreMLRequest的基础框架搭起来了,要实现多对象同时检测其实只要调整几个关键地方就行,我给你一步步拆解:
首先,换个支持多对象检测的模型
你现在用的Resnet50是图像分类模型,只能识别整张图的主类别,没法同时检测多个对象。得换成目标检测类的CoreML模型,比如YOLO系列、MobileNet SSD,或者Apple官方适配的YOLOv8n这类模型。
然后,优化代码结构并正确配置请求
先把模型初始化的逻辑移到帧捕获外面(别每次抓帧都重新创建模型,太耗性能),再补全VNCoreMLRequest的配置,添加结果回调来解析多对象数据:
extension CameraVC : AVCaptureVideoDataOutputSampleBufferDelegate { // 懒加载初始化检测模型,只执行一次 private lazy var detectionModel: VNCoreMLModel? = { do { // 替换成你的目标检测模型,比如YOLOv8n let coreMLModel = try YOLOv8n(configuration: MLModelConfiguration()).model return try VNCoreMLModel(for: coreMLModel) } catch { print("模型初始化失败:\(error.localizedDescription)") return nil } }() func captureOutput(_ output: AVCaptureOutput, didOutput sampleBuffer: CMSampleBuffer, from connection: AVCaptureConnection) { print("Camera was able to capture a frame: ", Date()) guard let pixelBuffer = CMSampleBufferGetImageBuffer(sampleBuffer) else { return } guard let model = detectionModel else { return } // 初始化VNCoreMLRequest,设置结果回调 let detectionRequest = VNCoreMLRequest(model: model) { [weak self] request, error in guard let self = self else { return } if let error = error { print("检测请求出错:\(error.localizedDescription)") return } // 解析多对象检测结果 guard let results = request.results as? [VNRecognizedObjectObservation] else { print("无法解析检测结果") return } // 遍历所有检测到的对象 for object in results { // 过滤低置信度结果(比如只保留置信度>0.5的) guard let topLabel = object.labels.first, topLabel.confidence > 0.5 else { continue } print("检测到:\(topLabel.identifier),置信度:\(String(format: "%.2f", topLabel.confidence))") print("位置(归一化坐标):\(object.boundingBox)") // 如果要在预览画面画框,必须回到主线程操作 DispatchQueue.main.async { // 这里可以添加绘制检测框的逻辑 self.drawBoundingBox(for: object) } } } // 设置图像缩放选项,适配模型输入要求 detectionRequest.imageCropAndScaleOption = .scaleFill // 用全局队列执行检测,避免阻塞相机回调 let imageHandler = VNImageRequestHandler(cvPixelBuffer: pixelBuffer, options: [:]) DispatchQueue.global(qos: .userInitiated).async { do { try imageHandler.perform([detectionRequest]) } catch { print("执行检测请求失败:\(error.localizedDescription)") } } } // 可选:添加绘制检测框的方法 private func drawBoundingBox(for observation: VNRecognizedObjectObservation) { let previewSize = previewView.bounds.size // 转换归一化坐标到UIKit坐标系(CoreML的原点在左下角,UIKit在左上角) let x = observation.boundingBox.origin.x * previewSize.width let y = (1 - observation.boundingBox.origin.y - observation.boundingBox.size.height) * previewSize.height let width = observation.boundingBox.size.width * previewSize.width let height = observation.boundingBox.size.height * previewSize.height let boundingBox = CGRect(x: x, y: y, width: width, height: height) // 这里可以创建CAShapeLayer来绘制矩形框,添加到previewView的layer上 let boxLayer = CAShapeLayer() boxLayer.path = UIBezierPath(rect: boundingBox).cgPath boxLayer.strokeColor = UIColor.red.cgColor boxLayer.fillColor = UIColor.clear.cgColor boxLayer.lineWidth = 2.0 previewView.layer.addSublayer(boxLayer) // 记得要移除旧的图层,避免叠加太多 // 可以在每次画框前先清除previewView上的所有boxLayer } }
最后,几个关键注意事项
- 模型转换:如果你的模型是PyTorch/TensorFlow格式,需要用CoreML Tools转换成CoreML格式,确保输出是
VNRecognizedObjectObservation类型。 - 性能优化:用懒加载初始化模型、把检测任务放到全局队列、设置置信度阈值过滤无效结果,这些都能提升实时检测的流畅度。
- 坐标系转换:CoreML返回的boundingBox是归一化坐标,且原点在左下角,和UIKit的坐标系不一样,必须转换后才能正确画框。
内容的提问来源于stack exchange,提问作者Barkley
相关产品推荐
相关产品推荐

