Swift中如何将[[Int]]传入MTLBuffer并在Metal端正确接收处理
实现目标
在Swift中定义[[Int]]类型的二维整型数组,通过MTLBuffer传入Metal,利用Metal多线程并行能力完成矩阵加法运算后,将结果返回给Swift侧,实现过程中遇到异常。
遇到的问题
Metal运行时提示graph不是合法指针,初步判断向Metal传参的方式存在错误,相关实现代码如下:
Swift侧代码
import MetalKit let graph: [[Int]] = [ [0, 1, 2, 999, 999, 999], [1, 0, 999, 5, 1, 999], [2, 999, 0, 2, 3, 999], [999, 5, 2, 0, 2, 2], [999, 1, 3, 2, 0, 1], [999, 999, 999, 2, 1, 0]] func fooFunc(gra: [[Int]]) { let count = gra.count let device = MTLCreateSystemDefaultDevice() let commandQueue = device?.makeCommandQueue() let gpuFunctionLibrary = device?.makeDefaultLibrary() let funcGPUFunction = gpuFunctionLibrary?.makeFunction(name: "MetalFunc") var funcPipelineState: MTLComputePipelineState! do { funcPipelineState = try device?.makeComputePipelineState(function: funcGPUFunction!) } catch { print(error) } let graphBuff = device?.makeBuffer(bytes: gra, length: MemoryLayout<Int>.size * count * count, options: .storageModeShared) let resultBuff = device?.makeBuffer(length: MemoryLayout<Int>.size * count, options: .storageModeShared) let commandBuffer = commandQueue?.makeCommandBuffer() let commandEncoder = commandBuffer?.makeComputeCommandEncoder() commandEncoder?.setComputePipelineState(additionComputePipelineState) commandEncoder?.setBuffer(graphBuff, offset: 0, index: 0) commandEncoder?.setBuffer(resultBuff, offset: 0, index: 1) let threadsPerGrid = MTLSize(width: count, height: 1, depth: 1) let maxThreadsPerThreadgroup = additionComputePipelineState.maxTotalThreadsPerThreadgroup // 1024 let threadsPerThreadgroup = MTLSize(width: maxThreadsPerThreadgroup, height: 1, depth: 1) commandEncoder?.dispatchThreads(threadsPerGrid, threadsPerThreadgroup: threadsPerThreadgroup) commandEncoder?.endEncoding() commandBuffer?.commit() commandBuffer?.waitUntilCompleted() let resultBufferPointer = resultBuff?.contents().bindMemory(to: Int.self, capacity: MemoryLayout<Int>.size * count) print("Result: \(Int(resultBufferPointer!.pointee) as Any)") } gpuDijkstra(gra: graph)
Metal侧代码
#include <metal_stdlib> using namespace metal; kernel void MetalFunc(constant int *graph [[ buffer(0) ]], constant int *result [[ buffer(1) ]], uint index [[ thread_position_in_grid ]] { const int size = sizeof(*graph); int result[size][size]; for(int k = 0; k<size; k++){ result[index][k]=graph[index][k]+graph[index][k]; //ERROR: Subscripted value is not an array, pointer, or vector } }
问题原因与修正方案
代码存在4个核心错误,逐一修正即可:
- Swift侧二维数组内存不连续:
[[Int]]是嵌套的Array结构体,每一行都是独立的内存块,整体不是连续的整型存储,直接传入gra作为buffer内容,GPU拿到的是Swift侧Array的内部结构指针,不是实际矩阵数据,访问必然非法。 - 缓冲区长度与类型不匹配:结果缓冲区只分配了
count个Int的长度,矩阵加法需要count*count个元素;同时Swift侧Int是64位,Metal侧int是32位,两边类型长度不对齐会出现数据错位。 - 代码变量名不一致:创建的计算管线状态叫
funcPipelineState,后续调用时用了不存在的additionComputePipelineState;定义的函数叫fooFunc,最后调用的是gpuDijkstra,会直接触发编译错误。 - Metal内核逻辑错误:
sizeof(*graph)计算的是单个int的长度(固定为4),根本不是矩阵边长;Metal(与C一致)不支持用运行时变量定义栈数组长度,int result[size][size]写法非法;graph是int*一维指针,不能直接用二维下标访问。
具体修正步骤
Swift侧调整
- 先把二维数组铺平为连续内存的一维数组,统一使用
Int32类型和Metal侧对齐 - 修正缓冲区长度,补传矩阵边长参数
- 统一变量名,读取结果时遍历所有元素
修正后核心代码:
import MetalKit let graph: [[Int32]] = [ [0, 1, 2, 999, 999, 999], [1, 0, 999, 5, 1, 999], [2, 999, 0, 2, 3, 999], [999, 5, 2, 0, 2, 2], [999, 1, 3, 2, 0, 1], [999, 999, 999, 2, 1, 0]] func metalMatrixAdd(gra: [[Int32]]) { let count = gra.count let totalCount = count * count let intSize = MemoryLayout<Int32>.size guard let device = MTLCreateSystemDefaultDevice(), let commandQueue = device.makeCommandQueue(), let gpuFunctionLibrary = device.makeDefaultLibrary(), let funcGPUFunction = gpuFunctionLibrary.makeFunction(name: "MetalFunc"), let funcPipelineState = try? device.makeComputePipelineState(function: funcGPUFunction) else { print("Metal initialization failed") return } // 铺平二维数组为连续一维内存 let flatGraph = gra.flatMap { $0 } guard let graphBuff = device.makeBuffer(bytes: flatGraph, length: intSize * totalCount, options: .storageModeShared), let resultBuff = device.makeBuffer(length: intSize * totalCount, options: .storageModeShared) else { print("Buffer creation failed") return } guard let commandBuffer = commandQueue.makeCommandBuffer(), let commandEncoder = commandBuffer.makeComputeCommandEncoder() else { print("Command creation failed") return } commandEncoder.setComputePipelineState(funcPipelineState) commandEncoder.setBuffer(graphBuff, offset: 0, index: 0) commandEncoder.setBuffer(resultBuff, offset: 0, index: 1) // 传矩阵边长 var matrixSize = Int32(count) commandEncoder.setBytes(&matrixSize, length: intSize, index: 2) let threadsPerGrid = MTLSize(width: count, height: 1, depth: 1) let maxThreadsPerThreadgroup = funcPipelineState.maxTotalThreadsPerThreadgroup let threadsPerThreadgroup = MTLSize(width: min(maxThreadsPerThreadgroup, count), height: 1, depth: 1) commandEncoder.dispatchThreads(threadsPerGrid, threadsPerThreadgroup: threadsPerThreadgroup) commandEncoder.endEncoding() commandBuffer.commit() commandBuffer.waitUntilCompleted() // 读取结果 let resultPointer = resultBuff.contents().bindMemory(to: Int32.self, capacity: totalCount) var result: [[Int32]] = [] for i in 0..<count { var row: [Int32] = [] for j in 0..<count { row.append(resultPointer[i * count + j]) } result.append(row) } print("Result: \(result)") } metalMatrixAdd(gra: graph)
Metal侧调整
- 新增矩阵边长参数
- 一维指针手动计算偏移访问元素,结果直接写入输出缓冲区,不需要定义栈上二维数组
修正后内核代码:
#include <metal_stdlib> using namespace metal; kernel void MetalFunc(constant int *graph [[ buffer(0) ]], device int *result [[ buffer(1) ]], constant int &matrixSize [[ buffer(2) ]], uint index [[ thread_position_in_grid ]]) { for(int k = 0; k < matrixSize; k++){ int offset = index * matrixSize + k; result[offset] = graph[offset] + graph[offset]; } }
内容的提问来源于stack exchange,提问作者blkpingu
相关产品推荐
相关产品推荐

