You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用CuPy结合cuda_fp16编译半精度CUDA内核遇编译错误求助

问题

我尝试用CuPy结合cuda_fp16头文件提供的半精度格式编译一个简单CUDA内核,内核代码如下:

code = r'''
extern "C" {

#include <cuda_fp16.h>

__global__ void kernel(half * const f1, half * const f2)
{
   if (blockDim.x*blockIdx.x + threadIdx.x < 12 && blockDim.y*blockIdx.y + threadIdx.y < 12)
   {
      const int ctr_0 = blockDim.x*blockIdx.x + threadIdx.x;
      const int ctr_1 = blockDim.y*blockIdx.y + threadIdx.y;
      f1[12*ctr_1 + ctr_0] = f2[12*ctr_1 + ctr_0];
   } 
}

}
'''

编译代码如下:

options = ('-I/path/to/cuda/include/', )

mod = cp.RawModule(code=code, options=options, backend="nvrtc", jitify=True)
func = mod.get_function("kernel")

编译后出现大量类似如下的错误:

cuda_fp16.hpp(266): error: more than one instance of overloaded function "operator++" has "C" linkage

cuda_fp16.hpp(267): error: more than one instance of overloaded function "operator--" has "C" linkage

...

共检测到24个编译错误

当前环境:cupy-cuda11x + CUDA 11.2

解决方案

问题出在你把<cuda_fp16.h>头文件放在了extern "C"块内部。C语言不支持函数重载,但cuda_fp16.h里定义了大量重载运算符,这些是C++特性,放到C链接块里必然会触发编译错误。

只需要把头文件包含移到extern "C"外面,只将内核函数放在extern "C"块中即可,修改后的内核代码:

code = r'''
#include <cuda_fp16.h>

extern "C" {

__global__ void kernel(half * const f1, half * const f2)
{
   if (blockDim.x*blockIdx.x + threadIdx.x < 12 && blockDim.y*blockIdx.y + threadIdx.y < 12)
   {
      const int ctr_0 = blockDim.x*blockIdx.x + threadIdx.x;
      const int ctr_1 = blockDim.y*blockIdx.y + threadIdx.y;
      f1[12*ctr_1 + ctr_0] = f2[12*ctr_1 + ctr_0];
   } 
}

}
'''

这样修改后,头文件的C++特性会正常编译,内核函数通过extern "C"保证链接兼容性,和CuPy的调用逻辑不冲突。

另外,新版本CuPy会自动配置CUDA的include路径,你可以尝试去掉options参数简化编译:

mod = cp.RawModule(code=code, backend="nvrtc", jitify=True)
func = mod.get_function("kernel")

内容的提问来源于stack exchange,提问作者Markus Holzer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 07:30:34