CoreML Swift开发中如何将图片列表转为MLMultiArray传入模型
解决方案
你需要先将每张UIImage调整为224×224的目标尺寸,提取RGB三通道的像素值,按照模型训练时的归一化规则处理后,填充到MLMultiArray的对应位置即可,具体实现如下:
1. 新增UIImage预处理扩展
添加两个工具方法,分别实现图片缩放、RGB像素值提取:
extension UIImage { /// 缩放图片到指定尺寸 func scaled(to size: CGSize) -> UIImage? { UIGraphicsBeginImageContextWithOptions(size, false, 1.0) defer { UIGraphicsEndImageContext() } draw(in: CGRect(origin: .zero, size: size)) return UIGraphicsGetImageFromCurrentImageContext() } /// 提取RGB三通道像素值,返回顺序为[R0,G0,B0,R1,G1,B1,...]的Float数组 func rgbPixelValues(normalizeClosure: (Float, Float, Float) -> (Float, Float, Float)) -> [Float]? { guard let cgImage = self.cgImage else { return nil } let width = cgImage.width let height = cgImage.height let bytesPerPixel = 4 let bytesPerRow = bytesPerPixel * width let pixelBuffer = UnsafeMutableRawPointer.allocate(byteCount: height * bytesPerRow, alignment: MemoryLayout<UInt8>.alignment) defer { pixelBuffer.deallocate() } let colorSpace = CGColorSpaceCreateDeviceRGB() guard let context = CGContext( data: pixelBuffer, width: width, height: height, bitsPerComponent: 8, bytesPerRow: bytesPerRow, space: colorSpace, bitmapInfo: CGImageAlphaInfo.premultipliedLast.rawValue ) else { return nil } context.draw(cgImage, in: CGRect(x: 0, y: 0, width: width, height: height)) var pixels = [Float]() pixels.reserveCapacity(width * height * 3) for y in 0..<height { for x in 0..<width { let offset = y * bytesPerRow + x * bytesPerPixel let r = Float(pixelBuffer.load(fromByteOffset: offset, as: UInt8.self)) let g = Float(pixelBuffer.load(fromByteOffset: offset+1, as: UInt8.self)) let b = Float(pixelBuffer.load(fromByteOffset: offset+2, as: UInt8.self)) let (normR, normG, normB) = normalizeClosure(r, g, b) pixels.append(normR) pixels.append(normG) pixels.append(normB) } } return pixels } }
2. 完善业务逻辑代码
替换你原有的循环逻辑,按如下方式填充MLMultiArray并调用模型推理:
let fileList = FileManager.default.enumerator(at: url, includingPropertiesForKeys: keys) let model = try Classification(configuration: MLModelConfiguration()) let targetSize = CGSize(width: 224, height: 224) // 归一化规则:请和你模型训练时的预处理规则保持一致,以下为两种常见示例 // 示例1:简单归一化到0~1区间 let normalize: (Float, Float, Float) -> (Float, Float, Float) = { r, g, b in (r / 255.0, g / 255.0, b / 255.0) } // 示例2:ImageNet预训练模型常用归一化 // let normalize: (Float, Float, Float) -> (Float, Float, Float) = { r, g, b in // ((r - 123.68) / 58.393, (g - 116.779) / 57.12, (b - 103.939) / 57.375) // } for case let file as URL in fileList { guard url.startAccessingSecurityScopedResource() else { continue } defer { url.stopAccessingSecurityScopedResource() } // 1. 加载并缩放图片 guard let originalImage = UIImage(contentsOfFile: file.path), let scaledImage = originalImage.scaled(to: targetSize), let pixelValues = scaledImage.rgbPixelValues(normalizeClosure: normalize) else { continue } // 2. 初始化当前图片对应的MLMultiArray guard let inputArray = try? MLMultiArray(shape: [1, 224, 224, 3], dataType: .float32) else { continue } // 3. 填充像素值到MLMultiArray for i in 0..<224 { for j in 0..<224 { for k in 0..<3 { let pixelIndex = i * 224 * 3 + j * 3 + k inputArray[[0, i as NSNumber, j as NSNumber, k as NSNumber]] = NSNumber(value: pixelValues[pixelIndex]) } } } // 4. 调用模型推理 do { let predictionResult = try model.prediction(input: inputArray) // 处理输出结果 print("图片\(file.lastPathComponent)推理完成") } catch { print("推理失败:\(error.localizedDescription)") } }
注意事项
- 如果你需要批量推理多张图片,可以把MLMultiArray的第一维修改为批量大小,按上述填充逻辑依次填充多张图片的像素值即可
- 若推理结果异常,可优先检查两个点:一是通道顺序是否正确(部分模型要求BGR通道,调整rgbPixelValues中返回r/g/b的顺序即可),二是归一化规则是否和训练时一致
- 生产环境建议去掉所有强制解包逻辑,添加完整的错误捕获处理
内容的提问来源于stack exchange,提问作者Yash Johri
相关产品推荐
相关产品推荐

