You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用DeepSpeed微调OPT 1.3B模型时遇运算符不匹配错误求助

问题:DeepSpeed微调OPT 1.3B时的CUDA算子编译错误

相关CUDA代码片段

template <typename T>
__global__ void moe_res_matmul(T* residual, T* coef, T* mlp_out, int seq_len, int hidden_dim)
{
    constexpr int granularity = 16;
    constexpr int vals_per_access = granularity / sizeof(T);

    T* residual_seq = residual + blockIdx.x * hidden_dim;
    T* mlp_out_seq = mlp_out + blockIdx.x * hidden_dim;

    for (unsigned tid = threadIdx.x * vals_per_access; tid < hidden_dim;
         tid += blockDim.x * vals_per_access) {
        T mlp[vals_per_access];
        T res[vals_per_access];
        T coef1[vals_per_access];
        T coef2[vals_per_access];

        mem_access::load_global<granularity>(mlp, mlp_out_seq + tid);
        mem_access::load_global<granularity>(res, residual_seq + tid);
        mem_access::load_global<granularity>(coef1, coef + tid);
        mem_access::load_global<granularity>(coef2, coef + tid + hidden_dim);

#pragma unroll
        for (int idx = 0; idx < vals_per_access; idx++) {
            mlp[idx] = mlp[idx] * coef2[idx] + res[idx] * coef1[idx];
        }

        mem_access::store_global<granularity>(mlp_out_seq + tid, mlp);
    }
}

错误日志

/.../python3.10/site-packages/deepspeed/ops/csrc/transformer/inference/csrc/gelu.cu(529): 
error: no operator "*" matches these operands
    operand types are: __half * __half
        mlp[idx] = mlp[idx] * coef2[idx] + res[idx] * coef1[idx];
                                  ^
    detected during:
        instantiation of "void moe_res_matmul(T *, T *, T *, int, int) [with T=__half]"
at line 547
        instantiation of "void launch_moe_res_matmul(T *, T *, T *, int, int, cudaStream_t) [with T=__half]"
at line 566

环境依赖

datasets>=2.8.0
sentencepiece>=0.1.97
protobuf==3.20.3
accelerate>=0.15.0
torch>=1.12.0
deepspeed>=0.9.0

解决思路

  • 匹配CUDA与PyTorch版本:PyTorch 1.12.0官方推荐搭配CUDA 11.3或11.6,确保系统CUDA版本和PyTorch编译依赖的CUDA版本一致,版本不匹配会导致__half类型的运算算子无法正确重载。
  • 升级DeepSpeed到最新稳定版:0.9.0版本存在该半精度算子的bug,升级到0.12.x及以上版本,官方已修复该类类型运算的兼容性问题。
  • 手动修复CUDA算子代码:若无法升级DeepSpeed,找到报错的gelu.cu文件,将半精度乘法替换为CUDA内置函数:把mlp[idx] * coef2[idx]改为__hmul(mlp[idx], coef2[idx]),res[idx] * coef1[idx]改为__hmul(res[idx], coef1[idx]),__half类型需用专门的CUDA intrinsics函数完成运算。
  • 临时禁用半精度训练:调整训练脚本,关闭FP16混合精度(AMP),改用FP32精度训练,绕过该半精度算子的编译问题,但会增加显存占用。
  • 从源码重新编译DeepSpeed:卸载当前预编译的DeepSpeed,从源码克隆编译,确保编译时使用的CUDA版本与PyTorch一致,源码编译可解决预编译包的兼容性问题。

内容的提问来源于stack exchange,提问作者coderLMN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 01:43:10