You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用std::index_sequence实现OpenCL内核加速比计算的编译期循环?

实现方案

你要的编译期展开效果可以用std::index_sequence配合折叠表达式(C17及以上)或者初始化列表展开(C14兼容)实现,完整代码如下:

#include <utility>
#include <iostream>
#include <chrono>

// 你原有的runOpenCL模板实现不需要改动
template<int threadcount>
INLINE void runOpenCL()
{
    constexpr int itemsPerThread = (MATRIX_HEIGHT + threadcount - 1) / threadcount;
    clObjs.physicsKernel.setArg(2, threadcount);
    clObjs.physicsKernel.setArg(3, itemsPerThread);
    clObjs.queue.enqueueNDRangeKernel(clObjs.physicsKernel, cl::NullRange, cl::NDRange(threadcount), cl::NullRange);
    clObjs.queue.finish();
    glFinish();
    const std::vector<cl::Memory> glObjs = { clObjs.glBuffer };
    clObjs.queue.enqueueAcquireGLObjects(&glObjs);
    clObjs.queue.enqueueCopyBuffer(clObjs.outBuffer, clObjs.glBuffer, 0, 0, planets_size_points);
    clObjs.queue.enqueueReleaseGLObjects(&glObjs);
}

// 编译期展开辅助函数
template <std::size_t... Is>
void runAllOpenCL(std::index_sequence<Is...>)
{
    // C++17折叠表达式写法,编译期逐个展开所有索引的调用
    (([&](){
        auto start = std::chrono::high_resolution_clock::now();
        runOpenCL<static_cast<int>(Is + 1)>(); // Is从0开始,加1对应你要的1起始计数
        auto end = std::chrono::high_resolution_clock::now();
        std::cout << std::chrono::duration<double, std::milli>(end - start).count() << "ms\n";
    }()), ...);
}

// C++14兼容版本,替换上面的辅助函数即可
/*
template <std::size_t... Is>
void runAllOpenCL(std::index_sequence<Is...>)
{
    int dummy[] = {0, ([&](){
        auto start = std::chrono::high_resolution_clock::now();
        runOpenCL<static_cast<int>(Is + 1)>();
        auto end = std::chrono::high_resolution_clock::now();
        std::cout << std::chrono::duration<double, std::milli>(end - start).count() << "ms\n";
    }(), 0)...};
    (void)dummy; // 消除未使用变量警告
}
*/

需要调用的时候只要一行代码即可完成1到99999的全部展开:

runAllOpenCL(std::make_index_sequence<99999>{});

注意:10万级别的模板实例化会显著增加编译时间、增大生成的二进制体积,实际测试时可以根据需求调整展开上限。

内容的提问来源于stack exchange,提问作者Jack Benson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 17:54:00