DeepLabV3模型MLMultiArray输出转CIImage速度过慢解决方案咨询
MLMultiArray转CIImage速度极慢的高效解决方案
问题背景
当前使用苹果官方提供的DeepLabV3 MLModel模型,模型预测过程速度很快:将连续输入的CIImage快速转换为CVPixelBuffer送入模型完成预测,但预测完成后需要将输出的分割掩码转换为CIImage时,即使在iPhone 13 Pro设备上该操作的运行速度也极慢。
获取预测结果的核心代码如下:
let prediction = try! deepLab.prediction(image: pixelBuffer) let semanticPredictions = prediction.semanticPredictions
已尝试两种获取分割掩码对应CIImage的方案,运行速度均达不到预期:
- 使用UIGraphicsImageRenderer逐像素绘制生成图像,对应实现代码:
func fetchMasKUsingGraphicRenderer(mlMultiArray: MLMultiArray) -> CIImage { let aWidth = CGFloat(mlMultiArray.shape[0].intValue) let aHeight = CGFloat(mlMultiArray.shape[1].intValue) let renderer = UIGraphicsImageRenderer(size: CGSize(width: aWidth, height: aHeight)) let img = renderer.image(actions: { context in let ctx = context.cgContext ctx.clear(CGRect(x: 0.0, y: 0.0, width: Double(aWidth), height: Double(aHeight))); for j in 0..<Int(aHeight) { for i in 0..<Int(aWidth) { let aValue = (mlMultiArray[j * Int(aHeight) + i].floatValue > 0.0) ? 1.0 : 0.0 let aRect = CGRect( x: CGFloat(i), y: CGFloat(j), width: 1.0, height: 1.0) let aColor: UIColor = UIColor( displayP3Red: 0.0, green: 0.0, blue: 0.0, alpha: CGFloat(aValue)) aColor.setFill() UIRectFill(aRect) } } }) return CIImage(image: img)! }
- 使用CoreMLHelpers工具库进行转换,核心调用代码:
let maskedCiImage = CIImage(cgImage: semanticPredictions.cgImage()!)
卡顿核心原因
两种方案慢的根源都是在CPU侧做逐像素遍历+跨内存域拷贝:逐像素绘制、CoreMLHelpers的cgImage()方法本质都是把MLMultiArray里的每个像素值挨个读取后写入CGImage对应内存,DeepLabV3输出分辨率为513*513时需要遍历26万+个像素,还要走CPU到GPU的内存传输,必然产生高延迟。
高效实现方案
最高效的方式是完全跳过CPU逐像素操作,直接基于Metal能力从MLMultiArray绑定的内存生成CIImage,全程走GPU实现零拷贝处理,具体实现如下:
基础转换实现
直接给MLMultiArray的内存加CVPixelBuffer包装,不做额外数据拷贝,二值化等后处理全部用GPU侧的Core Image滤镜完成:
import CoreML import CoreImage func mlMultiArrayToMaskCIImage(_ multiArray: MLMultiArray, threshold: Float = 0.0) -> CIImage? { // 校验DeepLabV3输出为单通道Float32格式 guard multiArray.dataType == .float32 else { return nil } let height = multiArray.shape[0].intValue let width = multiArray.shape[1].intValue // 创建和MLMultiArray共享内存的CVPixelBuffer,无拷贝开销 var pixelBuffer: CVPixelBuffer? let bufferAttrs = [ kCVPixelBufferCGImageCompatibilityKey: kCFBooleanFalse, kCVPixelBufferCGBitmapContextCompatibilityKey: kCFBooleanFalse, kCVPixelBufferMetalCompatibilityKey: kCFBooleanTrue ] as CFDictionary CVPixelBufferCreateWithBytes( kCFAllocatorDefault, width, height, kCVPixelFormatType_32Float, // 和MLMultiArray的Float32类型对齐 multiArray.dataPointer, multiArray.strides[0].intValue * MemoryLayout<Float32>.stride, nil, nil, bufferAttrs, &pixelBuffer ) guard let rawBuffer = pixelBuffer else { return nil } var maskImage = CIImage(cvPixelBuffer: rawBuffer) // iOS15+ 直接用内置阈值滤镜做二值化,全程GPU执行 if #available(iOS 15.0, *) { let thresholdFilter = CIFilter.colorThreshold() thresholdFilter.inputImage = maskImage thresholdFilter.threshold = threshold maskImage = thresholdFilter.outputImage ?? maskImage } else { // 低版本用自定义CIKernel实现阈值逻辑,同样运行在GPU guard let thresholdKernel = CIKernel(source: """ kernel vec4 thresholdKernel(sampler image, float threshold) { float pixelVal = sample(image, samplerCoord(image)).r; float alpha = pixelVal > threshold ? 1.0 : 0.0; return vec4(0.0, 0.0, 0.0, alpha); } """) else { return maskImage } maskImage = thresholdKernel.apply( extent: maskImage.extent, roiCallback: { _, rect in return rect }, arguments: [maskImage, threshold] ) ?? maskImage } return maskImage }
可选优化:生成可直接叠加的彩色掩码
如果需要把掩码叠加到原图上,无需手动绘制颜色,直接用Core Image内置滤镜调整通道即可:
// 将单通道掩码转为黑色透明、白色不透明的遮罩层 let colorMatrix = CIFilter.colorMatrix() colorMatrix.inputImage = maskImage colorMatrix.rVector = CIVector(x: 0, y: 0, z: 0, w: 0) colorMatrix.gVector = CIVector(x: 0, y: 0, z: 0, w: 0) colorMatrix.bVector = CIVector(x: 0, y: 0, z: 0, w: 0) colorMatrix.aVector = CIVector(x: 0, y: 0, z: 0, w: 1) colorMatrix.biasVector = CIVector(x: 0, y: 0, z: 0, w: 0) let finalOverlayMask = colorMatrix.outputImage
性能注意事项
- 禁止把MLMultiArray转成CGImage/UIImage后再转CIImage,这一步会强制触发CPU侧内存拷贝和格式转换,开销极大
- 所有图像变换(缩放、旋转、和原图叠加)都用CIFilter在GPU侧完成,不要在CPU侧做逐像素操作
- 可以在MLModel生成阶段直接配置输出格式为CVPixelBuffer,连MLMultiArray到CVPixelBuffer的包装步骤都能省略,速度还能进一步提升
CVPixelBufferCreateWithBytes创建的PixelBuffer和MLMultiArray共享同一块内存,不要在MLMultiArray释放后继续使用对应CIImage,需要长期持有图像时调用maskImage.clampedToExtent().renderedToCVPixelBuffer()拷贝一份独立内存即可
内容的提问来源于stack exchange,提问作者Roi Mulia
相关产品推荐
相关产品推荐

