You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CUDA库函数执行时间测量Timer类失效问题排查

CUDA性能计时异常问题解决

问题背景

开发CUDA库时,需要对比CPU与GPU函数的性能差异,编写了Timer类测量执行时间,但实际运行后得到的计时结果是极小的垃圾值,且所有m值对应的结果完全相同,而实际运行耗时有数分钟,执行时间理应随m增大而增加。

原Timer类代码

class Timer
{
public:
    Timer()
    {
        _StartTimepoint = std::chrono::steady_clock::now();
    }
 
    ~Timer() {}

    void Stop()
    {

        _stopped = true;
        using namespace std::chrono;
        auto endTimepoint = steady_clock::now();

        auto start = time_point_cast<milliseconds>(_StartTimepoint).time_since_epoch().count();
        auto end = time_point_cast<milliseconds>(endTimepoint).time_since_epoch().count();

        auto _ms = end - start;



        _secs   = _ms   / 1000;
        _ms    -= _secs * 1000;
        _mins   = _secs / 60;
        _secs  -= _mins * 60;
        _hour   = _mins / 60;
        _mins  -= _hour * 60;


    }


    double GetTime(){
        if(_stopped == true)
            return _ms;
        else{
            Stop();
            return _ms;
        }
    }

private:
    std::chrono::time_point< std::chrono::steady_clock> _StartTimepoint;
    double _secs,_ms,_mins,_hour;
    bool _stopped = false;
};

原测试循环代码

for (size_t m = MIN_M; m < MAX_M; m+=M_STEP){
        m_array[m_cont] = m;
        //simulate
        double time_gpu,time_cpu;

        Timer timer_gpu;
        run_device(prcr_args,seeds,&m_array[m_cont]);
        timer_gpu.Stop();
        time_gpu = timer_gpu.GetTime();

        Timer timer_cpu;
        simulate_host(prcr_args,seeds,&m_array[m_cont]);
        timer_cpu.Stop();
        time_cpu = timer_cpu.GetTime();
        
        double g = time_cpu/time_gpu;
        
        ofs << m  //stream to print the results
            << "," << time_cpu
            << "," << time_gpu 
            << "," << g << "\n";
        m_cont ++;
    }

异常输出

m,cpu_time,gpu_time,g
10,9.88131e-324,6.90979e-310,1.43004e-14
15,9.88131e-324,6.90979e-310,1.43004e-14
....
90,9.88131e-324,6.90979e-310,1.43004e-14
95,9.88131e-324,6.90979e-310,1.43004e-14
100,9.88131e-324,6.90979e-310,1.43004e-14

问题分析

  1. Timer类变量覆盖与未初始化问题

    • 在Stop()函数中,定义了局部变量auto _ms = end - start;,直接覆盖了类成员变量_ms,后续时间计算仅针对局部变量,类成员_ms始终未被赋值,处于未初始化状态,导致GetTime()返回内存垃圾值。
    • 类成员_secs、_ms、_mins、_hour在构造时未初始化,进一步触发未定义行为。
  2. CUDA异步执行导致计时不准确

    • CUDA核函数是异步执行的,CPU调用run_device(假设内部启动核函数)后会立即返回,不会等待GPU完成计算。此时timer_gpu.Stop()记录的只是核函数启动时间,而非实际执行完成时间。

修复方案

1. 修复Timer类

  • 移除局部变量_ms,直接使用类成员变量;
  • 构造函数中初始化所有时间成员变量;
  • 简化时间计算逻辑:
#include <chrono>

class Timer
{
public:
    Timer()
        : _StartTimepoint(std::chrono::steady_clock::now()),
          _secs(0.0), _ms(0.0), _mins(0.0), _hour(0.0),
          _stopped(false)
    {}
 
    ~Timer() {}

    void Stop()
    {
        if (_stopped) return; // 避免重复调用
        _stopped = true;
        using namespace std::chrono;
        auto endTimepoint = steady_clock::now();

        // 直接计算总毫秒数
        auto duration = duration_cast<milliseconds>(endTimepoint - _StartTimepoint);
        _ms = duration.count();
    }

    double GetTime()
    {
        if (!_stopped)
        {
            Stop();
        }
        return _ms;
    }

private:
    std::chrono::time_point<std::chrono::steady_clock> _StartTimepoint;
    double _secs, _ms, _mins, _hour;
    bool _stopped;
};

2. 处理CUDA异步执行问题

在调用run_device后添加CUDA同步函数,等待GPU完成计算后再停止计时:

for (size_t m = MIN_M; m < MAX_M; m+=M_STEP){
        m_array[m_cont] = m;
        double time_gpu,time_cpu;

        // GPU计时:添加同步等待
        Timer timer_gpu;
        run_device(prcr_args,seeds,&m_array[m_cont]);
        cudaError_t err = cudaDeviceSynchronize(); // 等待GPU完成
        if (err != cudaSuccess) {
            printf("CUDA sync error: %s\n", cudaGetErrorString(err));
            return 1;
        }
        timer_gpu.Stop();
        time_gpu = timer_gpu.GetTime();

        // CPU计时无需额外处理
        Timer timer_cpu;
        simulate_host(prcr_args,seeds,&m_array[m_cont]);
        timer_cpu.Stop();
        time_cpu = timer_cpu.GetTime();
        
        double g = time_cpu/time_gpu;
        
        ofs << m 
            << "," << time_cpu
            << "," << time_gpu 
            << "," << g << "\n";
        m_cont ++;
    }

额外建议

  • 对同一m值多次运行函数并取平均,减少偶然因素影响;
  • 在run_device调用后检查cudaGetLastError(),排查内部CUDA错误;
  • 若需要更高精度,可将duration_cast<milliseconds>改为duration_cast<microseconds>,并调整返回值格式。

内容的提问来源于stack exchange,提问作者magmar1968

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 21:55:20