You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用CUDA Thrust实现分段归约并存储部分结果?

用CUDA Thrust单次调用实现分段归约

要实现单次Thrust调用完成分段归约,最直接的方案是用thrust::reduce_by_key——这是Thrust专门用来按键分组归约的接口,完全匹配你的需求。

核心思路

构造一个键映射规则,让每个分块内的所有元素对应同一个键(比如分块大小为2时,索引0、1对应键0,索引2、3对应键1,以此类推),然后通过reduce_by_key把相同键的元素归约求和,最终直接得到每个分块的部分和数组。

完整代码示例

#include <thrust/device_vector.h>
#include <thrust/reduce.h>
#include <thrust/iterator/counting_iterator.h>
#include <thrust/iterator/transform_iterator.h>
#include <thrust/iterator/discard_iterator.h>
#include <iostream>

int main() {
    // 输入数据
    int data[] = {10,20,30,40,50,60,70,80};
    const int n = sizeof(data)/sizeof(data[0]);
    const int chunk_size = 2;
    const int num_chunks = n / chunk_size;

    // 拷贝数据到设备向量
    thrust::device_vector<int> d_data(data, data + n);
    thrust::device_vector<int> partial_sums(num_chunks);

    // 单次调用reduce_by_key完成分段求和
    thrust::reduce_by_key(
        // 生成键的起始迭代器:用索引映射成分块键
        thrust::make_transform_iterator(thrust::make_counting_iterator(0),
                                        [chunk_size] __device__ (int i) { return i / chunk_size; }),
        // 生成键的结束迭代器
        thrust::make_transform_iterator(thrust::make_counting_iterator(n),
                                        [chunk_size] __device__ (int i) { return i / chunk_size; }),
        // 待归约的原始数据
        d_data.begin(),
        // 丢弃输出的键(我们只需要部分和结果)
        thrust::discard_iterator(),
        // 部分和的输出数组
        partial_sums.begin(),
        // 键的比较规则:判断是否属于同一块
        thrust::equal_to<int>(),
        // 归约操作:求和
        thrust::plus<int>()
    );

    // 打印结果
    std::cout << "分段部分和数组:";
    for (int sum : partial_sums) {
        std::cout << sum << " ";
    }
    std::cout << std::endl;

    return 0;
}

代码说明

  • 用thrust::counting_iterator生成连续索引,再通过transform_iterator将索引映射为分块键(整数除法i / chunk_size自动完成分块分组),无需额外存储键数组,节省内存。
  • thrust::discard_iterator()用来忽略reduce_by_key输出的键数组,只保留我们需要的部分和结果。
  • 全程仅一次Thrust调用,充分利用GPU的并行计算能力,避免了循环方案的串行瓶颈。

对比循环方案

你之前的循环方案会多次调用thrust::reduce,每次仅处理一个分块,无法发挥GPU的并行优势。而reduce_by_key是批量并行处理所有分块,在数据量较大时性能提升非常明显,完全符合Thrust的并行编程理念。

内容的提问来源于stack exchange,提问作者Sangjun Lee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 03:08:10