咨询CUDA中返回自定义Vector类型函数的GPU执行方法
CUDA向量库函数GPU执行方案
明确__global__与__device__的定位
__global__函数是主机端发起调用的设备入口函数,强制要求返回void,且不能作为运算符重载(运算符重载用于表达式计算,和__global__作为核函数入口的定位不匹配)。用__device__函数实现向量运算逻辑
所有返回Vector<N,T>的运算(比如运算符+)都应该用__device__修饰,这类函数可以在设备端的核函数或其他__device__函数中调用,支持自定义类型返回,只要你的Vector类型是符合设备端要求的聚合类型(无主机端特有的成员或逻辑)。示例代码:
template<int N, typename T> struct Vector { T data[N]; // 设备端运算符重载 __device__ Vector<N, T> operator+(const Vector<N, T>& other) const { Vector<N, T> res; for (int i = 0; i < N; ++i) { res.data[i] = data[i] + other.data[i]; } return res; } };通过__global__核函数作为执行入口
要让这些__device__函数在GPU上运行,必须写一个__global__核函数作为主机到设备的入口,在核函数内部调用你的__device__向量运算函数。比如处理批量向量相加的核函数:template<int N, typename T> __global__ void vectorAddKernel(Vector<N,T>* inputA, Vector<N,T>* inputB, Vector<N,T>* output, int count) { int idx = blockIdx.x * blockDim.x + threadIdx.x; if (idx < count) { output[idx] = inputA[idx] + inputB[idx]; // 调用__device__运算符+ } }主机端调用流程
在主机端,你需要完成设备内存分配、数据拷贝、核函数启动、结果回收的完整流程,示例:int main() { const int count = 1024; const int vecSize = 3; // N=3 using Vec3f = Vector<3, float>; // 主机内存分配与初始化 Vec3f* h_a = new Vec3f[count]; Vec3f* h_b = new Vec3f[count]; Vec3f* h_out = new Vec3f[count]; // ... 初始化h_a和h_b的具体数据 // 设备内存分配 Vec3f* d_a; Vec3f* d_b; Vec3f* d_out; cudaMalloc(&d_a, count * sizeof(Vec3f)); cudaMalloc(&d_b, count * sizeof(Vec3f)); cudaMalloc(&d_out, count * sizeof(Vec3f)); // 数据拷贝到设备 cudaMemcpy(d_a, h_a, count * sizeof(Vec3f), cudaMemcpyHostToDevice); cudaMemcpy(d_b, h_b, count * sizeof(Vec3f), cudaMemcpyHostToDevice); // 启动核函数 int blockSize = 256; int gridSize = (count + blockSize - 1) / blockSize; vectorAddKernel<3, float><<<gridSize, blockSize>>>(d_a, d_b, d_out, count); cudaDeviceSynchronize(); // 等待核函数执行完成 // 结果拷贝回主机 cudaMemcpy(h_out, d_out, count * sizeof(Vec3f), cudaMemcpyDeviceToHost); // 清理内存 delete[] h_a; delete[] h_b; delete[] h_out; cudaFree(d_a); cudaFree(d_b); cudaFree(d_out); return 0; }排查__device__函数无效的可能原因
如果之前用__device__修饰后无效,大概率是你直接从主机端调用了__device__函数——这是不允许的,__device__函数只能在设备端执行(核函数内部或其他__device__函数中),必须通过__global__核函数间接调用。
内容的提问来源于stack exchange,提问作者merovingian
相关产品推荐
相关产品推荐

