You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DeepLabV3模型MLMultiArray输出转CIImage速度过慢解决方案咨询

MLMultiArray转CIImage速度极慢的高效解决方案

问题背景

当前使用苹果官方提供的DeepLabV3 MLModel模型,模型预测过程速度很快:将连续输入的CIImage快速转换为CVPixelBuffer送入模型完成预测,但预测完成后需要将输出的分割掩码转换为CIImage时,即使在iPhone 13 Pro设备上该操作的运行速度也极慢。

获取预测结果的核心代码如下:

let prediction = try! deepLab.prediction(image: pixelBuffer)
let semanticPredictions = prediction.semanticPredictions

已尝试两种获取分割掩码对应CIImage的方案,运行速度均达不到预期:

  • 使用UIGraphicsImageRenderer逐像素绘制生成图像,对应实现代码:
func fetchMasKUsingGraphicRenderer(mlMultiArray: MLMultiArray) -> CIImage {

  let aWidth = CGFloat(mlMultiArray.shape[0].intValue)
  let aHeight = CGFloat(mlMultiArray.shape[1].intValue)

  let renderer = UIGraphicsImageRenderer(size: CGSize(width: aWidth, height: aHeight))

  let img = renderer.image(actions: { context in
      let ctx = context.cgContext
      ctx.clear(CGRect(x: 0.0, y: 0.0, width: Double(aWidth), height: Double(aHeight)));
      for j in 0..<Int(aHeight) {
          for i in 0..<Int(aWidth) {

              let aValue =
                  (mlMultiArray[j * Int(aHeight) + i].floatValue > 0.0) ?
                      1.0 : 0.0
              let aRect = CGRect(
                  x: CGFloat(i),
                  y: CGFloat(j),
                  width: 1.0,
                  height: 1.0)

              let aColor: UIColor = UIColor(
                  displayP3Red: 0.0,
                  green: 0.0,
                  blue: 0.0,
                  alpha: CGFloat(aValue))

              aColor.setFill()
              UIRectFill(aRect)
          }
      }
  })

  return CIImage(image: img)!
}
  • 使用CoreMLHelpers工具库进行转换,核心调用代码:
let maskedCiImage = CIImage(cgImage: semanticPredictions.cgImage()!)

卡顿核心原因

两种方案慢的根源都是在CPU侧做逐像素遍历+跨内存域拷贝:逐像素绘制、CoreMLHelpers的cgImage()方法本质都是把MLMultiArray里的每个像素值挨个读取后写入CGImage对应内存,DeepLabV3输出分辨率为513*513时需要遍历26万+个像素,还要走CPU到GPU的内存传输,必然产生高延迟。

高效实现方案

最高效的方式是完全跳过CPU逐像素操作,直接基于Metal能力从MLMultiArray绑定的内存生成CIImage,全程走GPU实现零拷贝处理,具体实现如下:

基础转换实现

直接给MLMultiArray的内存加CVPixelBuffer包装,不做额外数据拷贝,二值化等后处理全部用GPU侧的Core Image滤镜完成:

import CoreML
import CoreImage

func mlMultiArrayToMaskCIImage(_ multiArray: MLMultiArray, threshold: Float = 0.0) -> CIImage? {
    // 校验DeepLabV3输出为单通道Float32格式
    guard multiArray.dataType == .float32 else { return nil }
    let height = multiArray.shape[0].intValue
    let width = multiArray.shape[1].intValue
    
    // 创建和MLMultiArray共享内存的CVPixelBuffer,无拷贝开销
    var pixelBuffer: CVPixelBuffer?
    let bufferAttrs = [
        kCVPixelBufferCGImageCompatibilityKey: kCFBooleanFalse,
        kCVPixelBufferCGBitmapContextCompatibilityKey: kCFBooleanFalse,
        kCVPixelBufferMetalCompatibilityKey: kCFBooleanTrue
    ] as CFDictionary
    
    CVPixelBufferCreateWithBytes(
        kCFAllocatorDefault,
        width,
        height,
        kCVPixelFormatType_32Float, // 和MLMultiArray的Float32类型对齐
        multiArray.dataPointer,
        multiArray.strides[0].intValue * MemoryLayout<Float32>.stride,
        nil,
        nil,
        bufferAttrs,
        &pixelBuffer
    )
    
    guard let rawBuffer = pixelBuffer else { return nil }
    var maskImage = CIImage(cvPixelBuffer: rawBuffer)
    
    // iOS15+ 直接用内置阈值滤镜做二值化,全程GPU执行
    if #available(iOS 15.0, *) {
        let thresholdFilter = CIFilter.colorThreshold()
        thresholdFilter.inputImage = maskImage
        thresholdFilter.threshold = threshold
        maskImage = thresholdFilter.outputImage ?? maskImage
    } else {
        // 低版本用自定义CIKernel实现阈值逻辑,同样运行在GPU
        guard let thresholdKernel = CIKernel(source: """
        kernel vec4 thresholdKernel(sampler image, float threshold) {
            float pixelVal = sample(image, samplerCoord(image)).r;
            float alpha = pixelVal > threshold ? 1.0 : 0.0;
            return vec4(0.0, 0.0, 0.0, alpha);
        }
        """) else { return maskImage }
        maskImage = thresholdKernel.apply(
            extent: maskImage.extent,
            roiCallback: { _, rect in return rect },
            arguments: [maskImage, threshold]
        ) ?? maskImage
    }
    
    return maskImage
}

可选优化:生成可直接叠加的彩色掩码

如果需要把掩码叠加到原图上,无需手动绘制颜色,直接用Core Image内置滤镜调整通道即可:

// 将单通道掩码转为黑色透明、白色不透明的遮罩层
let colorMatrix = CIFilter.colorMatrix()
colorMatrix.inputImage = maskImage
colorMatrix.rVector = CIVector(x: 0, y: 0, z: 0, w: 0)
colorMatrix.gVector = CIVector(x: 0, y: 0, z: 0, w: 0)
colorMatrix.bVector = CIVector(x: 0, y: 0, z: 0, w: 0)
colorMatrix.aVector = CIVector(x: 0, y: 0, z: 0, w: 1)
colorMatrix.biasVector = CIVector(x: 0, y: 0, z: 0, w: 0)
let finalOverlayMask = colorMatrix.outputImage

性能注意事项

  • 禁止把MLMultiArray转成CGImage/UIImage后再转CIImage,这一步会强制触发CPU侧内存拷贝和格式转换,开销极大
  • 所有图像变换(缩放、旋转、和原图叠加)都用CIFilter在GPU侧完成,不要在CPU侧做逐像素操作
  • 可以在MLModel生成阶段直接配置输出格式为CVPixelBuffer,连MLMultiArray到CVPixelBuffer的包装步骤都能省略,速度还能进一步提升
  • CVPixelBufferCreateWithBytes创建的PixelBuffer和MLMultiArray共享同一块内存,不要在MLMultiArray释放后继续使用对应CIImage,需要长期持有图像时调用maskImage.clampedToExtent().renderedToCVPixelBuffer()拷贝一份独立内存即可

内容的提问来源于stack exchange,提问作者Roi Mulia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.03 02:48:27