You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效传递纹理坐标至Metal顶点着色器实现带混合点图元绘制?

嘿,我来帮你解决这个顶点数量过多的性能问题!首先得明确你的核心需求:基于纹理每个像素的红色通道值绘制带混合的点,以此实现直方图累加效果,但1920×1080的纹理每秒30次绘制会产生200多万个顶点,确实会带来不小的性能开销。下面给你几个高效的优化方案,结合你的代码场景逐一说明:


方案1:用实例化渲染替代大量顶点,通过实例ID计算UV

你当前的代码里尝试了两种绘制方式,但其中使用instanceCount的调用存在一个关键错误:顶点着色器里用了vertex_id而非instance_id来区分不同像素。修正后,我们可以仅用1个顶点,通过实例ID计算每个像素的UV坐标,GPU处理实例化的效率远高于处理百万级顶点。

修改后的顶点着色器

vertex MappedVertex vertexShaderHistogramBlenderRed (
    texture2d<float, access::sample> inputTexture [[ texture(0) ]],
    unsigned int instanceId [[instance_id]] // 替换为实例ID
) {
    MappedVertex out;
    constexpr sampler s(s_address::clamp_to_edge, t_address::clamp_to_edge, min_filter::linear, mag_filter::linear, coord::pixel);
    ushort width = inputTexture.get_width();
    ushort height = inputTexture.get_height();
    
    // 通过实例ID计算当前像素的UV坐标
    float X = (instanceId % width) / float(width);
    float Y = (instanceId / width) / float(height);
    
    // 修正红色值的获取逻辑:将0-1的float值转为0-255的整数
    float redFloat = inputTexture.sample(s, float2(X,Y)).r;
    int red = int(redFloat * 255.0);
    
    // 计算点的位置:从屏幕左边缘(-1)到右边缘(1)均匀分布
    out.position = float4(-1.0 + (red * (2.0 / 255.0)), 0.0, 0.0, 1.0);
    out.pointSize = 1.0;
    out.colorFactor = half3(1.0, 0.0, 0.0);
    return out;
}

简化后的绘制代码

只保留实例化渲染的调用即可,顶点数从200多万降到1:

renderEncoder?.drawPrimitives(type: .point, vertexStart: 0, vertexCount: 1, instanceCount: metalTexture!.width*metalTexture!.height)

方案2:用Compute Shader预处理直方图,再用少量实例绘制点

这是性能最优的方案:先通过Compute Shader统计每个红色通道值的像素数量(直方图),再用256个实例(对应0-255的颜色区间)绘制点,彻底摆脱百万级顶点的开销。

步骤1:Compute Shader统计直方图

kernel void computeHistogram(
    texture2d<float, access::read> inputTexture [[ texture(0) ]],
    device uint* histogramBuffer [[ buffer(0) ]],
    uint2 gid [[ thread_position_in_grid ]]
) {
    ushort width = inputTexture.get_width();
    ushort height = inputTexture.get_height();
    
    // 确保线程在纹理范围内
    if (gid.x >= width || gid.y >= height) {
        return;
    }
    
    float2 uv = float2(gid.x / float(width), gid.y / float(height));
    constexpr sampler s(s_address::clamp_to_edge, t_address::clamp_to_edge, coord::pixel);
    float redFloat = inputTexture.sample(s, uv).r;
    int redBin = int(redFloat * 255.0);
    
    // 原子操作安全累加计数
    atomic_fetch_add_explicit(&histogramBuffer[redBin], 1, memory_order_relaxed);
}

步骤2:CPU端调度Compute Shader

// 创建直方图缓冲区(大小为256个uint)
let histogramBuffer = device.makeBuffer(length: 256 * MemoryLayout<UInt32>.stride, options: .storageModeShared)!

// 编码Compute命令
let computeEncoder = commandBuffer?.makeComputeCommandEncoder()
computeEncoder?.setComputePipelineState(histogramComputePipelineState!)
computeEncoder?.setTexture(metalTexture, index: 0)
computeEncoder?.setBuffer(histogramBuffer, offset: 0, index: 0)

// 设置线程组大小
let threadGroupSize = MTLSize(width: 16, height: 16, depth: 1)
let threadGroups = MTLSize(width: (metalTexture.width + threadGroupSize.width - 1)/threadGroupSize.width,
                           height: (metalTexture.height + threadGroupSize.height - 1)/threadGroupSize.height,
                           depth: 1)
computeEncoder?.dispatchThreadgroups(threadGroups, threadsPerThreadgroup: threadGroupSize)
computeEncoder?.endEncoding()

步骤3:绘制直方图点

顶点着色器根据实例ID(对应颜色区间)获取计数,输出点的位置和颜色:

vertex MappedVertex vertexShaderHistogramDrawer(
    device uint* histogramBuffer [[ buffer(0) ]],
    unsigned int instanceId [[instance_id]]
) {
    MappedVertex out;
    // 计算点的X位置:覆盖整个屏幕宽度
    float x = -1.0 + (instanceId * (2.0 / 255.0));
    out.position = float4(x, 0.0, 0.0, 1.0);
    out.pointSize = 1.0;
    
    // 根据计数计算颜色亮度(可根据需求调整比例)
    float count = float(histogramBuffer[instanceId]);
    float maxCount = float(metalTexture.width * metalTexture.height);
    float intensity = count / maxCount;
    out.colorFactor = half3(intensity, 0.0, 0.0);
    return out;
}

绘制代码只需256个实例:

renderEncoder?.setBuffer(histogramBuffer, offset: 0, index: 0)
renderEncoder?.drawPrimitives(type: .point, vertexStart: 0, vertexCount: 1, instanceCount: 256)

方案3:用全屏四边形替代点图元(性能最优)

如果不需要严格使用点图元,用一个全屏四边形+片段着色器的方式性能最好,仅需4个顶点即可完成直方图绘制:

全屏四边形顶点着色器

vertex float4 fullscreenQuadVertex(unsigned int vertexId [[vertex_id]]) {
    // 覆盖整个屏幕的顶点坐标
    const float2 positions[4] = {
        float2(-1.0, -1.0),
        float2(1.0, -1.0),
        float2(-1.0, 1.0),
        float2(1.0, 1.0)
    };
    return float4(positions[vertexId], 0.0, 1.0);
}

片段着色器

fragment half4 fragmentShaderHistogramDrawer(
    device uint* histogramBuffer [[ buffer(0) ]],
    float2 uv [[ texturecoord(0) ]] // UV范围0-1
) {
    // 根据X坐标计算对应的颜色区间
    int redBin = int(uv.x * 255.0);
    float count = float(histogramBuffer[redBin]);
    float maxCount = float(metalTexture.width * metalTexture.height);
    float intensity = count / maxCount;
    return half4(half3(intensity, 0.0, 0.0), 1.0);
}

绘制代码

renderEncoder?.drawPrimitives(type: .triangleStrip, vertexStart: 0, vertexCount: 4)

总结

  • 若坚持使用点图元:方案1(实例化渲染)能将顶点数降至1,显著降低开销;
  • 追求极致性能:方案2(Compute Shader预处理+少量实例)既保留点图元,又将绘制量从百万级降至256;
  • 无需点图元:方案3(全屏四边形)性能最优,适合大多数直方图展示场景。

内容的提问来源于stack exchange,提问作者Deepak Sharma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 10:11:13