Firefox Nightly中使用await mapAsync下载WebGPU buffer速度慢5000倍
问题:WebGPU Buffer回读性能在Firefox Nightly中异常缓慢
我需要将WebGPU Buffer下载到主机内存以进行CPU计算,但在Ubuntu 22.04系统的Firefox Nightly(128.0a1)中,该操作速度比预期慢5000倍——下载一个Buffer约需100毫秒,而等效的CuPy代码仅需0.02毫秒。相同代码在Chrome Unstable中每次迭代仅需约2毫秒,我怀疑这是Firefox Nightly的Bug,但仍想确认是否存在操作疏漏。
参考的CuPy代码
import cupy as cp import time buf = cp.array([1, 2, 3]) for _ in range(10): # Synchronize to make sure we are really measuring the correct time (not strictly necessary since .get() synchronizes implicitly) cp.cuda.Stream.null.synchronize() start_time = time.perf_counter() values = buf.get() cp.cuda.Stream.null.synchronize() elapsed_time = time.perf_counter() - start_time print(f"values: {values}, {elapsed_time * 1000:.6f} milliseconds")
WebGPU测试代码(JavaScript)
该代码执行以下操作:
- 从数组[1, 2, 3]创建WebGPU Buffer
- 将WebGPU Buffer复制到中间回读Buffer
- 测量将回读Buffer的值读取到values数组的耗时
async function main(){ if (!navigator.gpu) alert("WebGPU not supported"); const adapter = await navigator.gpu.requestAdapter(); if (!adapter) alert("WebGPU not supported"); const device = await adapter.requestDevice(); if (!device) alert("WebGPU not supported"); const values = new Float32Array([1, 2, 3]); // Create GPU buffer for values const buffer = device.createBuffer({ size: values.byteLength, usage: GPUBufferUsage.STORAGE | GPUBufferUsage.COPY_SRC | GPUBufferUsage.COPY_DST, mappedAtCreation: true, }); // Copy values to GPU buffer new values.constructor(buffer.getMappedRange()).set(values); buffer.unmap(); // Create readback buffer const readBuffer = device.createBuffer({ size: values.byteLength, usage: GPUBufferUsage.COPY_DST | GPUBufferUsage.MAP_READ, }); // Measure 10 times to make sure it was not a fluke for (var iteration = 0; iteration < 10; iteration++){ let startTime = window.performance.now(); // Copy buffer to readback buffer const commandEncoder = device.createCommandEncoder(); commandEncoder.copyBufferToBuffer(buffer, 0, readBuffer, 0, values.byteLength); device.queue.submit([commandEncoder.finish()]); // Wait for the GPU to finish await readBuffer.mapAsync(GPUMapMode.READ); // Clear values to make sure we are really reading from the buffer values.fill(0); // Copy values from readback buffer values.set(new values.constructor(readBuffer.getMappedRange())); // Unmap the readback buffer so we can use it again next iteration readBuffer.unmap(); // Measure elapsed time let elapsedTime = window.performance.now() - startTime; console.log("values", values, "time", elapsedTime, "milliseconds"); } } main();
可能的操作优化点(排查疏漏)
- 复用CommandEncoder:当前代码每次循环都创建新的
CommandEncoder,可以提前创建并复用,减少对象创建销毁的额外开销。 - 显式等待队列任务完成:在调用
mapAsync前,添加await device.queue.onSubmittedWorkDone(),确保测量的是纯粹的回读耗时,避免异步等待的额外延迟干扰。 - 指定高性能GPU:请求适配器时添加
powerPreference: "high-performance"参数,确保使用独立GPU而非集成显卡,集成显卡的回读性能通常更差。 - 移除冗余内存操作:每次循环中
values.fill(0)是测试用的冗余操作,可以移除,避免影响时间测量的准确性。
内容的提问来源于stack exchange,提问作者983
相关产品推荐
相关产品推荐

