You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Swift中如何将[[Int]]传入MTLBuffer并在Metal端正确接收处理

实现目标

在Swift中定义[[Int]]类型的二维整型数组,通过MTLBuffer传入Metal,利用Metal多线程并行能力完成矩阵加法运算后,将结果返回给Swift侧,实现过程中遇到异常。

遇到的问题

Metal运行时提示graph不是合法指针,初步判断向Metal传参的方式存在错误,相关实现代码如下:

Swift侧代码

import MetalKit

let graph: [[Int]] = [
    [0, 1, 2, 999, 999, 999],
    [1, 0, 999, 5, 1, 999],
    [2, 999, 0, 2, 3, 999],
    [999, 5, 2, 0, 2, 2],
    [999, 1, 3, 2, 0, 1],
    [999, 999, 999, 2, 1, 0]]

func fooFunc(gra: [[Int]]) {
    
    let count = gra.count

    let device = MTLCreateSystemDefaultDevice()

    let commandQueue = device?.makeCommandQueue()

    let gpuFunctionLibrary = device?.makeDefaultLibrary()

    let funcGPUFunction = gpuFunctionLibrary?.makeFunction(name: "MetalFunc")

    var funcPipelineState: MTLComputePipelineState!
    do {
        funcPipelineState = try device?.makeComputePipelineState(function: funcGPUFunction!)
    } catch {
      print(error)
    }

    let graphBuff = device?.makeBuffer(bytes: gra,
                                      length: MemoryLayout<Int>.size * count * count,
                                      options: .storageModeShared)
    let resultBuff = device?.makeBuffer(length: MemoryLayout<Int>.size * count,
                                        options: .storageModeShared)

    let commandBuffer = commandQueue?.makeCommandBuffer()

    let commandEncoder = commandBuffer?.makeComputeCommandEncoder()
    commandEncoder?.setComputePipelineState(additionComputePipelineState)

    commandEncoder?.setBuffer(graphBuff, offset: 0, index: 0)
    commandEncoder?.setBuffer(resultBuff, offset: 0, index: 1)

    let threadsPerGrid = MTLSize(width: count, height: 1, depth: 1)
    let maxThreadsPerThreadgroup = additionComputePipelineState.maxTotalThreadsPerThreadgroup // 1024
    let threadsPerThreadgroup = MTLSize(width: maxThreadsPerThreadgroup, height: 1, depth: 1)
    commandEncoder?.dispatchThreads(threadsPerGrid,
                                    threadsPerThreadgroup: threadsPerThreadgroup)

    commandEncoder?.endEncoding()
    commandBuffer?.commit()
    commandBuffer?.waitUntilCompleted()
    
    let resultBufferPointer = resultBuff?.contents().bindMemory(to: Int.self,
                                                                capacity: MemoryLayout<Int>.size * count)

    print("Result: \(Int(resultBufferPointer!.pointee) as Any)")
}


gpuDijkstra(gra: graph)

Metal侧代码

#include <metal_stdlib>
using namespace metal;

kernel void MetalFunc(constant int *graph        [[ buffer(0) ]],
                      constant int *result   [[ buffer(1) ]],
                      uint   index [[ thread_position_in_grid ]]
{


    const int size = sizeof(*graph);
    int result[size][size];

    for(int k = 0; k<size; k++){
        result[index][k]=graph[index][k]+graph[index][k]; //ERROR: Subscripted value is not an array, pointer, or vector
    }
}
问题原因与修正方案

代码存在4个核心错误,逐一修正即可:

  • Swift侧二维数组内存不连续:[[Int]]是嵌套的Array结构体,每一行都是独立的内存块,整体不是连续的整型存储,直接传入gra作为buffer内容,GPU拿到的是Swift侧Array的内部结构指针,不是实际矩阵数据,访问必然非法。
  • 缓冲区长度与类型不匹配:结果缓冲区只分配了count个Int的长度,矩阵加法需要count*count个元素;同时Swift侧Int是64位,Metal侧int是32位,两边类型长度不对齐会出现数据错位。
  • 代码变量名不一致:创建的计算管线状态叫funcPipelineState,后续调用时用了不存在的additionComputePipelineState;定义的函数叫fooFunc,最后调用的是gpuDijkstra,会直接触发编译错误。
  • Metal内核逻辑错误:sizeof(*graph)计算的是单个int的长度(固定为4),根本不是矩阵边长;Metal(与C一致)不支持用运行时变量定义栈数组长度,int result[size][size]写法非法;graph是int*一维指针,不能直接用二维下标访问。

具体修正步骤

Swift侧调整

  • 先把二维数组铺平为连续内存的一维数组,统一使用Int32类型和Metal侧对齐
  • 修正缓冲区长度,补传矩阵边长参数
  • 统一变量名,读取结果时遍历所有元素
    修正后核心代码:
import MetalKit

let graph: [[Int32]] = [
    [0, 1, 2, 999, 999, 999],
    [1, 0, 999, 5, 1, 999],
    [2, 999, 0, 2, 3, 999],
    [999, 5, 2, 0, 2, 2],
    [999, 1, 3, 2, 0, 1],
    [999, 999, 999, 2, 1, 0]]

func metalMatrixAdd(gra: [[Int32]]) {
    let count = gra.count
    let totalCount = count * count
    let intSize = MemoryLayout<Int32>.size
    
    guard let device = MTLCreateSystemDefaultDevice(),
          let commandQueue = device.makeCommandQueue(),
          let gpuFunctionLibrary = device.makeDefaultLibrary(),
          let funcGPUFunction = gpuFunctionLibrary.makeFunction(name: "MetalFunc"),
          let funcPipelineState = try? device.makeComputePipelineState(function: funcGPUFunction) else {
        print("Metal initialization failed")
        return
    }

    // 铺平二维数组为连续一维内存
    let flatGraph = gra.flatMap { $0 }
    guard let graphBuff = device.makeBuffer(bytes: flatGraph,
                                            length: intSize * totalCount,
                                            options: .storageModeShared),
          let resultBuff = device.makeBuffer(length: intSize * totalCount,
                                             options: .storageModeShared) else {
        print("Buffer creation failed")
        return
    }

    guard let commandBuffer = commandQueue.makeCommandBuffer(),
          let commandEncoder = commandBuffer.makeComputeCommandEncoder() else {
        print("Command creation failed")
        return
    }
    
    commandEncoder.setComputePipelineState(funcPipelineState)
    commandEncoder.setBuffer(graphBuff, offset: 0, index: 0)
    commandEncoder.setBuffer(resultBuff, offset: 0, index: 1)
    // 传矩阵边长
    var matrixSize = Int32(count)
    commandEncoder.setBytes(&matrixSize, length: intSize, index: 2)

    let threadsPerGrid = MTLSize(width: count, height: 1, depth: 1)
    let maxThreadsPerThreadgroup = funcPipelineState.maxTotalThreadsPerThreadgroup
    let threadsPerThreadgroup = MTLSize(width: min(maxThreadsPerThreadgroup, count), height: 1, depth: 1)
    commandEncoder.dispatchThreads(threadsPerGrid,
                                    threadsPerThreadgroup: threadsPerThreadgroup)

    commandEncoder.endEncoding()
    commandBuffer.commit()
    commandBuffer.waitUntilCompleted()
    
    // 读取结果
    let resultPointer = resultBuff.contents().bindMemory(to: Int32.self, capacity: totalCount)
    var result: [[Int32]] = []
    for i in 0..<count {
        var row: [Int32] = []
        for j in 0..<count {
            row.append(resultPointer[i * count + j])
        }
        result.append(row)
    }
    print("Result: \(result)")
}

metalMatrixAdd(gra: graph)

Metal侧调整

  • 新增矩阵边长参数
  • 一维指针手动计算偏移访问元素,结果直接写入输出缓冲区,不需要定义栈上二维数组
    修正后内核代码:
#include <metal_stdlib>
using namespace metal;

kernel void MetalFunc(constant int *graph      [[ buffer(0) ]],
                      device int *result        [[ buffer(1) ]],
                      constant int &matrixSize  [[ buffer(2) ]],
                      uint index [[ thread_position_in_grid ]])
{
    for(int k = 0; k < matrixSize; k++){
        int offset = index * matrixSize + k;
        result[offset] = graph[offset] + graph[offset];
    }
}

内容的提问来源于stack exchange,提问作者blkpingu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 14:24:25