You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CUDA Thrust高效实现device_vector复制与分块求和的方法咨询

Thrust内置函数实现向量重复复制与分块求和

1. 向量重复复制(将device_vectorA重复4次到device_vectorB)

不用循环调用复制接口,直接用thrust::transform结合索引映射完成单次设备端操作,相比多次copy_n调用效率更高:

#include <thrust/device_vector.h>
#include <thrust/transform.h>
#include <thrust/iterator/counting_iterator.h>

int main() {
    const int A_size = 5;
    const int repeat_times = 4;
    const int B_size = A_size * repeat_times;

    thrust::device_vector<int> A(A_size);
    thrust::sequence(A.begin(), A.end(), 1); // 初始化A为1-5

    thrust::device_vector<int> B(B_size);

    // 一次transform完成重复复制
    thrust::transform(
        thrust::counting_iterator<int>(0),
        thrust::counting_iterator<int>(B_size),
        B.begin(),
        [A_size] __device__(int i) {
            return A[i % A_size]; // 通过模运算映射到A的索引
        }
    );

    // 验证输出(可选)
    thrust::device_vector<int> host_B(B_size);
    thrust::copy(B.begin(), B.end(), host_B.begin());
    for (int val : host_B) printf("%d ", val);
    // 输出:1 2 3 4 5 1 2 3 4 5 1 2 3 4 5 1 2 3 4 5
    return 0;
}

原理:counting_iterator生成B的元素索引序列,transform对每个索引计算i % A_size,直接从A中取对应元素赋值给B,全程是单次设备kernel调用。

2. 分块求和(device_vectorB按每5个元素分块求和到device_vectorC)

使用Thrust专门的分组归约函数thrust::reduce_by_key,替代手动循环分块求和:

#include <thrust/device_vector.h>
#include <thrust/reduce.h>
#include <thrust/iterator/counting_iterator.h>
#include <thrust/functional.h>

int main() {
    const int block_size = 5;
    const int block_count = 4;
    const int B_size = block_size * block_count;

    thrust::device_vector<int> B(B_size);
    thrust::sequence(B.begin(), B.end(), 1); // 初始化B为1-20

    thrust::device_vector<int> keys(B_size);
    // 生成分组键:每个5元素块内的元素对应相同键(0,0,...1,1,...3,3)
    thrust::transform(
        thrust::counting_iterator<int>(0),
        thrust::counting_iterator<int>(B_size),
        keys.begin(),
        [block_size] __device__(int i) {
            return i / block_size;
        }
    );

    thrust::device_vector<int> C(block_count);
    // 执行分组求和
    thrust::reduce_by_key(
        keys.begin(), keys.end(),
        B.begin(),
        thrust::make_discard_iterator(), // 不需要保留输出的键
        C.begin(),
        thrust::equal_to<int>(),         // 键匹配规则
        thrust::plus<int>()              // 归约操作:求和
    );

    // 验证输出(可选)
    thrust::device_vector<int> host_C(block_count);
    thrust::copy(C.begin(), C.end(), host_C.begin());
    for (int val : host_C) printf("%d ", val);
    // 输出:15 40 65 90
    return 0;
}

原理:先通过i / block_size为每个元素生成分组键,reduce_by_key会自动将连续相同键的元素求和,最终得到每个块的结果,全程是单次优化后的设备归约操作。

内容的提问来源于stack exchange,提问作者Dylan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 15:25:14