You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何测量cuda::convolution函数中FFT与IFFT的执行耗时?

如何测量CUDA FFT/IFFT的执行耗时

刚好我之前也碰到过类似的需求,要精准测量CUDA里FFT和IFFT的耗时,用CPU的计时器肯定不行——毕竟GPU操作是异步的,得用CUDA专门的事件API来做才准确。下面给你一步步拆解实现方法,直接套到你的代码里就行。

核心思路:用CUDA事件捕获时间戳

CUDA提供了cudaEvent_t来记录GPU操作的开始和结束时间,它能精准同步GPU的执行状态,避免异步操作带来的计时误差。我们要做的就是在FFT/IFFT调用前后分别记录事件,然后计算两个事件的时间差。

具体实现步骤

1. 提前创建CUDA事件

在你的函数开头(比如初始化plan之后),创建两个用于FFT计时、两个用于IFFT计时的事件,记得要加错误检查:

cudaEvent_t start_fft, stop_fft, start_ifft, stop_ifft;
cudaSafeCall(cudaEventCreate(&start_fft));
cudaSafeCall(cudaEventCreate(&stop_fft));
cudaSafeCall(cudaEventCreate(&start_ifft));
cudaSafeCall(cudaEventCreate(&stop_ifft));

(这里的cudaSafeCall和你用的cufftSafeCall逻辑一致,是封装了错误检查的宏,确保事件创建成功)

2. 在FFT/IFFT前后插入事件记录

把事件插入到你的cufftExecR2C和cufftExecC2R调用前后,必须和你的流_stream关联,不然计时会和其他GPU操作混在一起,完全不准:

比如第一个模板的FFT:

// 记录FFT开始事件
cudaSafeCall(cudaEventRecord(start_fft, _stream));
cufftSafeCall( cufftExecR2C(planR2C, templ_block.ptr<cufftReal>(), templ_spect.ptr<cufftComplex>()) );
// 记录FFT结束事件
cudaSafeCall(cudaEventRecord(stop_fft, _stream));

然后循环里的每个块的FFT和IFFT:

// 定义累计耗时变量,统计所有块的总时间
float total_fft_time = 0.0f;
float total_ifft_time = 0.0f;
int block_count = 0;

for (int y = 0; y < result.rows; y += block_size.height) {
    for (int x = 0; x < result.cols; x += block_size.width) {
        block_count++;
        Size image_roi_size(std::min(x + dft_size.width, image.cols) - x, std::min(y + dft_size.height, image.rows) - y);
        GpuMat image_roi(image_roi_size, CV_32F, (void*)(image.ptr<float>(y) + x), image.step);
        cuda::copyMakeBorder(image_roi, image_block, 0, image_block.rows - image_roi.rows, 0, image_block.cols - image_roi.cols, 0, Scalar(), _stream);
        
        // 记录当前块FFT开始
        cudaSafeCall(cudaEventRecord(start_fft, _stream));
        cufftSafeCall(cufftExecR2C(planR2C, image_block.ptr<cufftReal>(), image_spect.ptr<cufftComplex>()));
        cudaSafeCall(cudaEventRecord(stop_fft, _stream));
        
        cuda::mulAndScaleSpectrums(image_spect, templ_spect, result_spect, 0, 1.f / dft_size.area(), ccorr, _stream);
        
        // 记录IFFT开始
        cudaSafeCall(cudaEventRecord(start_ifft, _stream));
        cufftSafeCall(cufftExecC2R(planC2R, result_spect.ptr<cufftComplex>(), result_data.ptr<cufftReal>()));
        cudaSafeCall(cudaEventRecord(stop_ifft, _stream));
        
        Size result_roi_size(std::min(x + block_size.width, result.cols) - x, std::min(y + block_size.height, result.rows) - y);
        GpuMat result_roi(result_roi_size, result.type(), (void*)(result.ptr<float>(y) + x), result.step);
        GpuMat result_block(result_roi_size, result_data.type(), result_data.ptr(), result_data.step);
        result_block.copyTo(result_roi, _stream);
        
        // 计算当前块的FFT耗时
        float fft_time_ms = 0.0f;
        cudaSafeCall(cudaEventSynchronize(stop_fft)); // 等待GPU完成该FFT操作
        cudaSafeCall(cudaEventElapsedTime(&fft_time_ms, start_fft, stop_fft));
        total_fft_time += fft_time_ms;
        
        // 计算当前块的IFFT耗时
        float ifft_time_ms = 0.0f;
        cudaSafeCall(cudaEventSynchronize(stop_ifft));
        cudaSafeCall(cudaEventElapsedTime(&ifft_time_ms, start_ifft, stop_ifft));
        total_ifft_time += ifft_time_ms;
    }
}

3. 输出统计结果

循环结束后,你可以输出总耗时和平均每块的耗时:

printf("所有块FFT总耗时: %.3f ms,平均每块: %.3f ms\n", total_fft_time, total_fft_time / block_count);
printf("所有块IFFT总耗时: %.3f ms,平均每块: %.3f ms\n", total_ifft_time, total_ifft_time / block_count);

4. 销毁事件

在函数末尾,记得销毁创建的事件,避免内存泄漏:

cudaSafeCall(cudaEventDestroy(start_fft));
cudaSafeCall(cudaEventDestroy(stop_fft));
cudaSafeCall(cudaEventDestroy(start_ifft));
cudaSafeCall(cudaEventDestroy(stop_ifft));

关键注意事项

  • 流的关联:一定要把事件和你的_stream绑定,否则如果用默认流,会和其他GPU操作混在一起,计时完全不准。
  • 事件同步:调用cudaEventSynchronize确保GPU操作完成后再读取时间差,不然得到的是未完成的无效时间。
  • 多次测量取平均:单次测量可能受系统负载影响,建议多次运行函数,取平均耗时,结果更可靠。
  • 错误检查:所有CUDA和CuFFT操作都要加错误检查(像你用的cufftSafeCall),如果计时异常,先排查是否有API调用失败的情况。

内容的提问来源于stack exchange,提问作者Zaw

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:22:29