You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用CUDA实现数组步长为2的相邻两元素配对相加?

Implementing Pairwise Adjacent Sum (Step 2) in CUDA

Hey there! Let's walk through how to pull off this pairwise sum operation in CUDA. It's a straightforward task once you align the problem with CUDA's thread-based parallelism model. Here's a breakdown of the approach, code examples, and key considerations:

Core Idea

Each CUDA thread will handle one pair of elements: thread 0 adds a[0] + a[1], thread 1 adds a[2] + a[3], and so on. This way, we leverage parallelism efficiently since each sum operation is independent of the others.

1. The CUDA Kernel Function

This is where the actual addition happens. Each thread calculates its global ID, then maps that to the corresponding pair in the input array.

__global__ void pairwiseSum(const float* input, float* output, int numElements) {
    // Calculate the global thread index
    int idx = blockIdx.x * blockDim.x + threadIdx.x;
    
    // Only process if both elements of the pair exist
    if (2 * idx + 1 < numElements) {
        output[idx] = input[2 * idx] + input[2 * idx + 1];
    }
    // Note: If the input array has an odd number of elements, the last element
    // won't be processed here—we'll handle that in the host code below
}

2. Host Code Setup

We need to manage memory allocation, data transfer between host and device, kernel launch, and result retrieval. Here's a complete example:

#include <iostream>
#include <vector>

int main() {
    // Example input array (can be any length, even or odd)
    std::vector<float> h_input = {1.0f, 2.0f, 3.0f, 4.0f, 5.0f, 6.0f, 7.0f};
    int numInputElements = h_input.size();
    // Calculate output size: for odd lengths, we include the unpaired last element
    int numOutputElements = (numInputElements + 1) / 2;

    // Allocate memory on the GPU
    float* d_input;
    float* d_output;
    cudaMalloc(&d_input, numInputElements * sizeof(float));
    cudaMalloc(&d_output, numOutputElements * sizeof(float));

    // Copy input data from host to GPU
    cudaMemcpy(d_input, h_input.data(), numInputElements * sizeof(float), cudaMemcpyHostToDevice);

    // Configure kernel launch parameters
    int blockSize = 256; // Typical choice for CUDA (multiple of 32)
    // Calculate grid size to cover all output elements (ceiling division)
    int gridSize = (numOutputElements + blockSize - 1) / blockSize;

    // Launch the kernel
    pairwiseSum<<<gridSize, blockSize>>>(d_input, d_output, numInputElements);
    cudaDeviceSynchronize(); // Wait for kernel execution to finish

    // Handle odd-length input: copy the last unpaired element to output
    if (numInputElements % 2 != 0) {
        cudaMemcpy(&d_output[numOutputElements - 1], &h_input[numInputElements - 1], 
                   sizeof(float), cudaMemcpyHostToDevice);
    }

    // Copy results back from GPU to host
    std::vector<float> h_output(numOutputElements);
    cudaMemcpy(h_output.data(), d_output, numOutputElements * sizeof(float), cudaMemcpyDeviceToHost);

    // Print results for verification
    std::cout << "Input array: ";
    for (float val : h_input) std::cout << val << " ";
    std::cout << "\nPairwise sum output: ";
    for (float val : h_output) std::cout << val << " ";
    std::cout << std::endl;

    // Clean up GPU memory
    cudaFree(d_input);
    cudaFree(d_output);

    return 0;
}

Key Considerations

  • Thread Mapping: Each thread's global index directly maps to the index of the sum pair in the output array. We use 2*idx and 2*idx+1 to access the two elements in the input pair.
  • Odd Length Handling: If your input array has an odd number of elements, you can either ignore the last element or copy it directly to the output (as shown above).
  • Launch Configuration: Using a block size of 256 (or another multiple of 32) optimizes performance for most CUDA devices. The grid size is calculated with ceiling division to ensure we have enough threads to cover all pairs.
  • Error Checking: For production code, add error checks for cudaMalloc, cudaMemcpy, and kernel launches (use cudaGetLastError() after kernel launch to catch issues).

Alternative: Use the Thrust Library

If you prefer a high-level approach without writing raw CUDA kernels, you can use the Thrust library (included with CUDA). Here's a quick example:

#include <thrust/device_vector.h>
#include <thrust/transform.h>
#include <thrust/iterator/zip_iterator.h>
#include <thrust/tuple.h>

// Functor to add two elements from a tuple
struct addPair {
    __host__ __device__
    float operator()(const thrust::tuple<float, float>& pair) {
        return thrust::get<0>(pair) + thrust::get<1>(pair);
    }
};

int main() {
    std::vector<float> h_input = {1.0f, 2.0f, 3.0f, 4.0f, 5.0f, 6.0f};
    thrust::device_vector<float> d_input(h_input.begin(), h_input.end());
    int numOutputElements = d_input.size() / 2;
    thrust::device_vector<float> d_output(numOutputElements);

    // Use zip_iterator to pair adjacent elements, then apply the sum
    thrust::transform(
        thrust::make_zip_iterator(thrust::make_tuple(d_input.begin(), d_input.begin() + 1)),
        thrust::make_zip_iterator(thrust::make_tuple(d_input.end() - 1, d_input.end())),
        d_output.begin(),
        addPair()
    );

    // Copy results back to host and print (same as before)
    return 0;
}

内容的提问来源于stack exchange,提问作者yuexia_wrp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:58:31