Windows 11 + CUDA 12.5环境下nvcc编译CUDA核函数报错:attribute "__global__" does not apply here
Windows 11 + CUDA 12.5环境下nvcc编译CUDA核函数报错:attribute "global" does not apply here
嗨,我来帮你排查这个问题!从报错信息来看,主要有两个核心问题:一是nvcc把CUDA的__global__关键字错误识别成了MSVC的扩展属性,二是uint类型未定义。咱们一步步解决:
1. 解决uint未定义的问题
Windows下的MSVC编译器默认不支持uint作为unsigned int的简写别名(这是C99的扩展特性,MSVC需要手动开启或改用标准类型)。你有两个简单的解决方向:
- 直接把代码里所有的
uint替换成unsigned int; - 在代码开头添加
<stdint.h>头文件,然后将uint替换成标准的uint32_t类型。
2. 修复__global__关键字被错误解析的问题
报错里出现__declspec(__global__)说明nvcc没有正确识别CUDA的核函数修饰符,大概率是编译器没按CUDA代码规则处理你的文件。试试这几个方案:
- 强制指定文件类型:在编译命令里加上
-x cu,明确告诉nvcc这是CUDA源文件:nvcc -arch=sm_89 -x cu .\simplest_kernel.cu - 检查文件后缀:确认你的文件确实是
.cu后缀,而不是不小心改成了.cpp或.c(虽然你说其他.cu文件能编译,但还是确认下这个文件的后缀是否正确); - 确认编译环境:确保你是从NVIDIA CUDA命令提示符打开的PowerShell,而不是普通的PowerShell——只有前者会正确配置CUDA相关的环境变量。
修正后的示例代码
把uint替换成unsigned int后的代码如下:
#include <cuda_runtime.h> #include <iostream> #include <vector> __global__ void kernel(unsigned int *A, unsigned int *B, int row) { auto x = threadIdx.x / 4; auto y = threadIdx.x % 4; A[x * row + y] = x; B[x * row + y] = y; } int main(int argc, char **argv) { unsigned int *Xs, *Ys; unsigned int *Xs_d, *Ys_d; unsigned int SIZE = 4; Xs = (unsigned int *)malloc(SIZE * SIZE * sizeof(unsigned int)); Ys = (unsigned int *)malloc(SIZE * SIZE * sizeof(unsigned int)); cudaMalloc((void **)&Xs_d, SIZE * SIZE * sizeof(unsigned int)); cudaMalloc((void **)&Ys_d, SIZE * SIZE * sizeof(unsigned int)); dim3 grid_size(1, 1, 1); dim3 block_size(4 * 4); kernel<<<grid_size, block_size>>>(Xs_d, Ys_d, 4); cudaMemcpy(Xs, Xs_d, SIZE * SIZE * sizeof(unsigned int), cudaMemcpyDeviceToHost); cudaMemcpy(Ys, Ys_d, SIZE * SIZE * sizeof(unsigned int), cudaMemcpyDeviceToHost); cudaDeviceSynchronize(); for (int row = 0; row < SIZE; ++row) { for (int col = 0; col < SIZE; ++col) { std::cout << "[" << Xs[row * SIZE + col] << "|" << Ys[row * SIZE + col] << "] "; } std::cout << "\n"; } cudaFree(Xs_d); cudaFree(Ys_d); free(Xs); free(Ys); }
备注:内容来源于stack exchange,提问作者He Huang
相关产品推荐
相关产品推荐

