如何用CUDA实现数组步长为2的相邻两元素配对相加?
Hey there! Let's walk through how to pull off this pairwise sum operation in CUDA. It's a straightforward task once you align the problem with CUDA's thread-based parallelism model. Here's a breakdown of the approach, code examples, and key considerations:
Core Idea
Each CUDA thread will handle one pair of elements: thread 0 adds a[0] + a[1], thread 1 adds a[2] + a[3], and so on. This way, we leverage parallelism efficiently since each sum operation is independent of the others.
1. The CUDA Kernel Function
This is where the actual addition happens. Each thread calculates its global ID, then maps that to the corresponding pair in the input array.
__global__ void pairwiseSum(const float* input, float* output, int numElements) { // Calculate the global thread index int idx = blockIdx.x * blockDim.x + threadIdx.x; // Only process if both elements of the pair exist if (2 * idx + 1 < numElements) { output[idx] = input[2 * idx] + input[2 * idx + 1]; } // Note: If the input array has an odd number of elements, the last element // won't be processed here—we'll handle that in the host code below }
2. Host Code Setup
We need to manage memory allocation, data transfer between host and device, kernel launch, and result retrieval. Here's a complete example:
#include <iostream> #include <vector> int main() { // Example input array (can be any length, even or odd) std::vector<float> h_input = {1.0f, 2.0f, 3.0f, 4.0f, 5.0f, 6.0f, 7.0f}; int numInputElements = h_input.size(); // Calculate output size: for odd lengths, we include the unpaired last element int numOutputElements = (numInputElements + 1) / 2; // Allocate memory on the GPU float* d_input; float* d_output; cudaMalloc(&d_input, numInputElements * sizeof(float)); cudaMalloc(&d_output, numOutputElements * sizeof(float)); // Copy input data from host to GPU cudaMemcpy(d_input, h_input.data(), numInputElements * sizeof(float), cudaMemcpyHostToDevice); // Configure kernel launch parameters int blockSize = 256; // Typical choice for CUDA (multiple of 32) // Calculate grid size to cover all output elements (ceiling division) int gridSize = (numOutputElements + blockSize - 1) / blockSize; // Launch the kernel pairwiseSum<<<gridSize, blockSize>>>(d_input, d_output, numInputElements); cudaDeviceSynchronize(); // Wait for kernel execution to finish // Handle odd-length input: copy the last unpaired element to output if (numInputElements % 2 != 0) { cudaMemcpy(&d_output[numOutputElements - 1], &h_input[numInputElements - 1], sizeof(float), cudaMemcpyHostToDevice); } // Copy results back from GPU to host std::vector<float> h_output(numOutputElements); cudaMemcpy(h_output.data(), d_output, numOutputElements * sizeof(float), cudaMemcpyDeviceToHost); // Print results for verification std::cout << "Input array: "; for (float val : h_input) std::cout << val << " "; std::cout << "\nPairwise sum output: "; for (float val : h_output) std::cout << val << " "; std::cout << std::endl; // Clean up GPU memory cudaFree(d_input); cudaFree(d_output); return 0; }
Key Considerations
- Thread Mapping: Each thread's global index directly maps to the index of the sum pair in the output array. We use
2*idxand2*idx+1to access the two elements in the input pair. - Odd Length Handling: If your input array has an odd number of elements, you can either ignore the last element or copy it directly to the output (as shown above).
- Launch Configuration: Using a block size of 256 (or another multiple of 32) optimizes performance for most CUDA devices. The grid size is calculated with ceiling division to ensure we have enough threads to cover all pairs.
- Error Checking: For production code, add error checks for
cudaMalloc,cudaMemcpy, and kernel launches (usecudaGetLastError()after kernel launch to catch issues).
Alternative: Use the Thrust Library
If you prefer a high-level approach without writing raw CUDA kernels, you can use the Thrust library (included with CUDA). Here's a quick example:
#include <thrust/device_vector.h> #include <thrust/transform.h> #include <thrust/iterator/zip_iterator.h> #include <thrust/tuple.h> // Functor to add two elements from a tuple struct addPair { __host__ __device__ float operator()(const thrust::tuple<float, float>& pair) { return thrust::get<0>(pair) + thrust::get<1>(pair); } }; int main() { std::vector<float> h_input = {1.0f, 2.0f, 3.0f, 4.0f, 5.0f, 6.0f}; thrust::device_vector<float> d_input(h_input.begin(), h_input.end()); int numOutputElements = d_input.size() / 2; thrust::device_vector<float> d_output(numOutputElements); // Use zip_iterator to pair adjacent elements, then apply the sum thrust::transform( thrust::make_zip_iterator(thrust::make_tuple(d_input.begin(), d_input.begin() + 1)), thrust::make_zip_iterator(thrust::make_tuple(d_input.end() - 1, d_input.end())), d_output.begin(), addPair() ); // Copy results back to host and print (same as before) return 0; }
内容的提问来源于stack exchange,提问作者yuexia_wrp

