You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CUDA核函数受CPU代码影响变慢的原因排查求助

CUDA实现性能突降问题分析请求

我用CUDA实现了一个掩码模板匹配算法,功能测试正常。但在通过以下代码对比该CUDA实现与OpenCV CPU实现时,出现异常情况:

for (int i = 0; i < 200; i++){
    // ---------------------------------------- GPU --------------------------------------------
    double t1 = (double)cv::getTickCount();
    gpu_img.upload(img);
    matcher->match(gpu_img, gpu_result);    // gpu
    gpu_result.download(result);
    t1 = ((double)cv::getTickCount() - t1) / cv::getTickFrequency();
    std::cout << t1 * 1000 << std::endl;

    double max_val_gpu;
    double min_val_gpu;
    cv::Point max_loc_gpu;
    cv::minMaxLoc(result, &min_val_gpu, &max_val_gpu, 0, &max_loc_gpu);
    std::cout << "min_val_gpu: " << min_val_gpu << "  max_val_gpu: " << max_val_gpu << std::endl;
    std::cout << "max_loc_gpu: " << max_loc_gpu.x << " " << max_loc_gpu.y << std::endl;
    // -------------------------------------------------------------------------------------------

    // ---------------------------------------- CPU -----------------------------------------------
    double t4 = (double)cv::getTickCount();
    cv::matchTemplate(img, temp, result, 5, mask);  // cpu
    t4 = ((double)cv::getTickCount() - t4) / cv::getTickFrequency();
    std::cout << "cpu version: " << t4 * 1000 << " ms" << std::endl;

    double min_val_cpu;
    double max_val_cpu;
    cv::Point max_loc_cpu;
    cv::minMaxLoc(result, &min_val_cpu, &max_val_cpu, 0, &max_loc_cpu);
    std::cout << "min_val_cpu: " << min_val_cpu << " max_val_cpu: " << max_val_cpu << std::endl;
    std::cout << "max_loc_cpu: " << max_loc_cpu.x << " " << max_loc_cpu.y << std::endl;
    // ---------------------------------------------------------------------------------------------
}

异常现象

CUDA函数的耗时在若干次迭代后突然从4ms增至20ms:
The time consumption of the cuda function

我测试了其他CUDA函数(包括简单向量加法及OpenCV CUDA API),发现所有CUDA函数都会受CPU代码影响变慢:

  • 若每次迭代运行两次CPU函数,CUDA核函数耗时会在约第55次迭代时突增(同样从4ms到20ms);
  • 即使将CPU代码替换为单个waitKey(100)语句,CUDA核函数仍会变慢。

运行环境

  • 系统:Win10
  • CUDA版本:11.1
  • 编译环境:VS2015 + nvcc 11.1
  • 显卡:RTX 3060

nsys性能分析结果

  • 带CPU代码时的性能剖面:
    with cpu code
  • 无CPU代码时的性能剖面:
    without cpu code

可见所有CUDA API及核函数均出现性能下降,请求分析原因。


内容的提问来源于stack exchange,提问作者song

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 08:57:37