You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CoreML Swift开发中如何将图片列表转为MLMultiArray传入模型

解决方案

你需要先将每张UIImage调整为224×224的目标尺寸,提取RGB三通道的像素值,按照模型训练时的归一化规则处理后,填充到MLMultiArray的对应位置即可,具体实现如下:


1. 新增UIImage预处理扩展

添加两个工具方法,分别实现图片缩放、RGB像素值提取:

extension UIImage {
    /// 缩放图片到指定尺寸
    func scaled(to size: CGSize) -> UIImage? {
        UIGraphicsBeginImageContextWithOptions(size, false, 1.0)
        defer { UIGraphicsEndImageContext() }
        draw(in: CGRect(origin: .zero, size: size))
        return UIGraphicsGetImageFromCurrentImageContext()
    }
    
    /// 提取RGB三通道像素值,返回顺序为[R0,G0,B0,R1,G1,B1,...]的Float数组
    func rgbPixelValues(normalizeClosure: (Float, Float, Float) -> (Float, Float, Float)) -> [Float]? {
        guard let cgImage = self.cgImage else { return nil }
        let width = cgImage.width
        let height = cgImage.height
        let bytesPerPixel = 4
        let bytesPerRow = bytesPerPixel * width
        let pixelBuffer = UnsafeMutableRawPointer.allocate(byteCount: height * bytesPerRow, alignment: MemoryLayout<UInt8>.alignment)
        defer { pixelBuffer.deallocate() }
        
        let colorSpace = CGColorSpaceCreateDeviceRGB()
        guard let context = CGContext(
            data: pixelBuffer,
            width: width,
            height: height,
            bitsPerComponent: 8,
            bytesPerRow: bytesPerRow,
            space: colorSpace,
            bitmapInfo: CGImageAlphaInfo.premultipliedLast.rawValue
        ) else { return nil }
        
        context.draw(cgImage, in: CGRect(x: 0, y: 0, width: width, height: height))
        var pixels = [Float]()
        pixels.reserveCapacity(width * height * 3)
        for y in 0..<height {
            for x in 0..<width {
                let offset = y * bytesPerRow + x * bytesPerPixel
                let r = Float(pixelBuffer.load(fromByteOffset: offset, as: UInt8.self))
                let g = Float(pixelBuffer.load(fromByteOffset: offset+1, as: UInt8.self))
                let b = Float(pixelBuffer.load(fromByteOffset: offset+2, as: UInt8.self))
                let (normR, normG, normB) = normalizeClosure(r, g, b)
                pixels.append(normR)
                pixels.append(normG)
                pixels.append(normB)
            }
        }
        return pixels
    }
}

2. 完善业务逻辑代码

替换你原有的循环逻辑,按如下方式填充MLMultiArray并调用模型推理:

let fileList = FileManager.default.enumerator(at: url, includingPropertiesForKeys: keys)
let model = try Classification(configuration: MLModelConfiguration())
let targetSize = CGSize(width: 224, height: 224)

// 归一化规则:请和你模型训练时的预处理规则保持一致,以下为两种常见示例
// 示例1:简单归一化到0~1区间
let normalize: (Float, Float, Float) -> (Float, Float, Float) = { r, g, b in
    (r / 255.0, g / 255.0, b / 255.0)
}
// 示例2:ImageNet预训练模型常用归一化
// let normalize: (Float, Float, Float) -> (Float, Float, Float) = { r, g, b in
//     ((r - 123.68) / 58.393, (g - 116.779) / 57.12, (b - 103.939) / 57.375)
// }

for case let file as URL in fileList {
    guard url.startAccessingSecurityScopedResource() else { continue }
    defer { url.stopAccessingSecurityScopedResource() }
    
    // 1. 加载并缩放图片
    guard let originalImage = UIImage(contentsOfFile: file.path),
          let scaledImage = originalImage.scaled(to: targetSize),
          let pixelValues = scaledImage.rgbPixelValues(normalizeClosure: normalize) else {
        continue
    }
    
    // 2. 初始化当前图片对应的MLMultiArray
    guard let inputArray = try? MLMultiArray(shape: [1, 224, 224, 3], dataType: .float32) else {
        continue
    }
    
    // 3. 填充像素值到MLMultiArray
    for i in 0..<224 {
        for j in 0..<224 {
            for k in 0..<3 {
                let pixelIndex = i * 224 * 3 + j * 3 + k
                inputArray[[0, i as NSNumber, j as NSNumber, k as NSNumber]] = NSNumber(value: pixelValues[pixelIndex])
            }
        }
    }
    
    // 4. 调用模型推理
    do {
        let predictionResult = try model.prediction(input: inputArray)
        // 处理输出结果
        print("图片\(file.lastPathComponent)推理完成")
    } catch {
        print("推理失败:\(error.localizedDescription)")
    }
}

注意事项

  • 如果你需要批量推理多张图片,可以把MLMultiArray的第一维修改为批量大小,按上述填充逻辑依次填充多张图片的像素值即可
  • 若推理结果异常,可优先检查两个点:一是通道顺序是否正确(部分模型要求BGR通道,调整rgbPixelValues中返回r/g/b的顺序即可),二是归一化规则是否和训练时一致
  • 生产环境建议去掉所有强制解包逻辑,添加完整的错误捕获处理

内容的提问来源于stack exchange,提问作者Yash Johri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 05:15:03