You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Windows 11 + CUDA 12.5环境下nvcc编译CUDA核函数报错:attribute "__global__" does not apply here

Windows 11 + CUDA 12.5环境下nvcc编译CUDA核函数报错:attribute "global" does not apply here

嗨,我来帮你排查这个问题!从报错信息来看,主要有两个核心问题:一是nvcc把CUDA的__global__关键字错误识别成了MSVC的扩展属性,二是uint类型未定义。咱们一步步解决:

1. 解决uint未定义的问题

Windows下的MSVC编译器默认不支持uint作为unsigned int的简写别名(这是C99的扩展特性,MSVC需要手动开启或改用标准类型)。你有两个简单的解决方向:

  • 直接把代码里所有的uint替换成unsigned int;
  • 在代码开头添加<stdint.h>头文件,然后将uint替换成标准的uint32_t类型。

2. 修复__global__关键字被错误解析的问题

报错里出现__declspec(__global__)说明nvcc没有正确识别CUDA的核函数修饰符,大概率是编译器没按CUDA代码规则处理你的文件。试试这几个方案:

  • 强制指定文件类型:在编译命令里加上-x cu,明确告诉nvcc这是CUDA源文件:
    nvcc -arch=sm_89 -x cu .\simplest_kernel.cu
    
  • 检查文件后缀:确认你的文件确实是.cu后缀,而不是不小心改成了.cpp或.c(虽然你说其他.cu文件能编译,但还是确认下这个文件的后缀是否正确);
  • 确认编译环境:确保你是从NVIDIA CUDA命令提示符打开的PowerShell,而不是普通的PowerShell——只有前者会正确配置CUDA相关的环境变量。

修正后的示例代码

把uint替换成unsigned int后的代码如下:

#include <cuda_runtime.h>
#include <iostream>
#include <vector>

__global__ void kernel(unsigned int *A, unsigned int *B, int row) {
  auto x = threadIdx.x / 4;
  auto y = threadIdx.x % 4;
  A[x * row + y] = x;
  B[x * row + y] = y;
}

int main(int argc, char **argv) {
  unsigned int *Xs, *Ys;
  unsigned int *Xs_d, *Ys_d;

  unsigned int SIZE = 4;

  Xs = (unsigned int *)malloc(SIZE * SIZE * sizeof(unsigned int));
  Ys = (unsigned int *)malloc(SIZE * SIZE * sizeof(unsigned int));

  cudaMalloc((void **)&Xs_d, SIZE * SIZE * sizeof(unsigned int));
  cudaMalloc((void **)&Ys_d, SIZE * SIZE * sizeof(unsigned int));

  dim3 grid_size(1, 1, 1);
  dim3 block_size(4 * 4);

  kernel<<<grid_size, block_size>>>(Xs_d, Ys_d, 4);

  cudaMemcpy(Xs, Xs_d, SIZE * SIZE * sizeof(unsigned int), cudaMemcpyDeviceToHost);
  cudaMemcpy(Ys, Ys_d, SIZE * SIZE * sizeof(unsigned int), cudaMemcpyDeviceToHost);

  cudaDeviceSynchronize();

  for (int row = 0; row < SIZE; ++row) {
    for (int col = 0; col < SIZE; ++col) {
      std::cout << "[" << Xs[row * SIZE + col] << "|" << Ys[row * SIZE + col]
                << "] ";
    }
    std::cout << "\n";
  }

  cudaFree(Xs_d);
  cudaFree(Ys_d);
  free(Xs);
  free(Ys);
}

备注:内容来源于stack exchange,提问作者He Huang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.15 10:23:06