计算着色器中快速累加RWStructuredBuffer<float>的最优实现问询
我在光线生成着色器rayGen中向RWStructuredBuffer<float> tau写入数据,DispatchRays维度为(tauWidth, tauHeight, 1),tau的元素数量恰好为tauWidth * tauHeight,每个rayGen实例写入唯一元素。每帧会按顺序执行rayGen和计算着色器resolve:
[numthreads(16, 16, 1)] void resolve(uint3 dispatchThreadId : SV_DispatchThreadID) { execute(dispatchThreadId.xy); }
其中execute函数以(frameWidth, frameHeight, 1)维度调用,需写入RWTexture2D out,且需要tau所有元素的总和tau_sum。
我希望找到让execute能最快获取tau_sum的方法,此前查阅资料时对多步归约及CPU回读累加的方案感到困惑。
补充实现代码
HLSL(实际为SLANG)代码:
RWStructuredBuffer<float> tau; cbuffer reduceTauCB { const uint stride, size; } [numthreads(256, 1, 1)] void reduceTau(const uint3 dispatchThreadID: SV_DispatchThreadID) { tau[stride * i] = tau[stride * i] + tau[stride * i + stride / 2]; } [numthreads(1, 1, 1)] void reduceTauFinalize(const uint3 dispatchThreadID: SV_DispatchThreadID) { for (uint i = 1; i < size; ++i) tau[0] += tau[i * stride]; }
C++(基于Falcor)代码:
static std::uint32_t constexpr groupSize = 16 * 16; std::uint32_t const tauCount = mStaticParams.tauWidth * mStaticParams.tauHeight; auto reduceTauPassRootVar = mpReduceTauPass->getRootVar()["reduceTauCB"]; std::uint32_t stride = 2, threadCount = tauCount / 2; while (threadCount >= 256) { reduceTauPassRootVar["stride"] = stride; mpReduceTauPass->execute(pRenderContext, threadCount, 1, 1); stride *= 2; threadCount /= 2; } reduceTauPassRootVar["stride"] = stride; reduceTauPassRootVar["size"] = threadCount; mpReduceTauFinalizePass->execute(pRenderContext, 1, 1, 1);
我想知道该实现是否为最优性能方案,以及是否必须使用cbuffer,是否有直接向计算着色器传递参数的方法?
1. 当前实现的性能分析
你的归约实现属于全局内存归约,并非最优方案,核心问题有两点:
- 全局内存访问效率低:每一步归约直接读写全局内存的
tau缓冲区,未利用线程组共享内存(groupshared)做局部归约。GPU共享内存带宽远高于全局内存,先将数据加载到共享内存完成组内归约,再写回全局内存,能大幅减少全局内存访问次数,提升性能。 - 单线程累加瓶颈:
reduceTauFinalize用单线程循环累加剩余元素,当剩余元素数量较多时,这一步会成为明显的性能瓶颈。
更优的归约流程应该是:
- 每个线程组加载一批全局内存数据到共享内存,完成组内归约后将每组结果写回全局内存。
- 重复上述步骤,直到全局内存中只剩一个线程组的结果。
- 最后用一个线程组完成最终累加,而非单线程。
另外,CPU回读累加的方案性能通常更差——GPU到CPU的数据传输延迟高,且CPU并行计算能力远不如GPU,除非tau元素数量极小,否则不建议采用。
2. 是否必须使用cbuffer?
不是必须的。在现代GPU API(如DX12、Vulkan)及Falcor框架中,有多种参数传递方式:
- 根常量(Root Constant):可以直接将标量参数作为根常量传递,无需打包到cbuffer。Falcor的
RootVar支持直接设置根常量,对于stride和size这类简单标量,这种方式更直接。 - 推送常量(Push Constants):类似根常量,是Vulkan中的概念,Falcor对DX12和Vulkan做了封装,也支持用这种方式传递小批量标量参数。
不过cbuffer也有优势:若需传递多个参数,打包到cbuffer可减少根参数数量,避免根签名过于复杂。
3. 直接向计算着色器传递参数的方法
以Falcor为例,可通过根常量直接传递参数,修改后的代码如下:
修改SLANG代码,使用根常量
RWStructuredBuffer<float> tau; // 直接定义根常量,无需cbuffer [[vk::push_constant]] [[dx12::root_constant]] const uint stride; [[vk::push_constant]] [[dx12::root_constant]] const uint size; [numthreads(256, 1, 1)] void reduceTau(const uint3 dispatchThreadID: SV_DispatchThreadID) { uint i = dispatchThreadID.x; tau[stride * i] = tau[stride * i] + tau[stride * i + stride / 2]; } [numthreads(1, 1, 1)] void reduceTauFinalize(const uint3 dispatchThreadID: SV_DispatchThreadID) { for (uint i = 1; i < size; ++i) tau[0] += tau[i * stride]; }
修改C++代码,设置根常量
static std::uint32_t constexpr groupSize = 16 * 16; std::uint32_t const tauCount = mStaticParams.tauWidth * mStaticParams.tauHeight; auto reduceTauPassRootVar = mpReduceTauPass->getRootVar(); std::uint32_t stride = 2, threadCount = tauCount / 2; while (threadCount >= 256) { // 直接设置根常量 reduceTauPassRootVar["stride"] = stride; mpReduceTauPass->execute(pRenderContext, threadCount, 1, 1); stride *= 2; threadCount /= 2; } // 设置最终化阶段的根常量 auto finalizeRootVar = mpReduceTauFinalizePass->getRootVar(); finalizeRootVar["stride"] = stride; finalizeRootVar["size"] = threadCount; mpReduceTauFinalizePass->execute(pRenderContext, 1, 1, 1);
此外,也可将参数打包到结构化缓冲区传递,但对于少量标量参数,根常量或推送常量效率更高——它们可直接被着色器读取,无需额外内存访问。
内容的提问来源于stack exchange,提问作者0xbadf00d

