You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在主机端对CUDA的__half类型执行算术运算?

在CUDA 11.2主机端操作__half的可行方案

1. 直接调用CUDA提供的主机端半精度运算函数

CUDA 11.2的cuda_fp16.hpp中已提供主机端可用的半精度运算函数,无需手动转float再转回。这些函数与设备端底层函数命名一致,直接调用即可:

  • 加法:__host__ __half __hadd(const __half a, const __half b)
  • 减法:__host__ __half __hsub(const __half a, const __half b)
  • 乘法:__host__ __half __hmul(const __half a, const __half b)
  • 除法:__host__ __half __hdiv(const __half a, const __half b)

使用示例:

#include <cuda_fp16.hpp>

int main() {
    __half a = __float2half(1.5f);
    __half b = __float2half(2.0f);
    __half sum = __hadd(a, b);       // 主机端直接计算半精度加法
    __half product = __hmul(a, b);   // 主机端直接计算半精度乘法
    return 0;
}

2. 重载主机端的__half算术运算符

如果想保持和设备端一致的运算符语法,可以自行在主机端重载__half的+、-、*、/运算符,内部调用对应的主机端运算函数:

#include <cuda_fp16.hpp>

// 主机端重载加法运算符
__host__ __forceinline__ __half operator+(const __half& lh, const __half& rh) {
    return __hadd(lh, rh);
}

// 主机端重载减法运算符
__host__ __forceinline__ __half operator-(const __half& lh, const __half& rh) {
    return __hsub(lh, rh);
}

// 主机端重载乘法运算符
__host__ __forceinline__ __half operator*(const __half& lh, const __half& rh) {
    return __hmul(lh, rh);
}

// 主机端重载除法运算符
__host__ __forceinline__ __half operator/(const __half& lh, const __half& rh) {
    return __hdiv(lh, rh);
}

int main() {
    __half a = __float2half(1.5f);
    __half b = __float2half(2.0f);
    __half sum = a + b;       // 直接用运算符,与设备端语法一致
    __half diff = a - b;
    __half product = a * b;
    __half quotient = a / b;
    return 0;
}

注意:重载时用__host__限定符,确保仅在主机端生效,避免和设备端的运算符定义冲突。

3. 使用CUDA C++ API的cuda::half类型

CUDA 11及以上版本提供的CUDA C++ API中,cuda::half类型原生支持主机端算术运算符重载,且与底层__half类型兼容(可通过类型转换或构造函数互相转换),是最便捷的方案。

使用示例:

#include <cuda/std/half.hpp>

int main() {
    cuda::half a(1.5f);
    cuda::half b(2.0f);
    cuda::half sum = a + b;       // 直接使用运算符
    cuda::half product = a * b;

    // 和__half互相转换
    __half raw_a = static_cast<__half>(a);
    cuda::half wrapped_b(raw_a);
    return 0;
}

内容的提问来源于stack exchange,提问作者einpoklum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 10:07:37