使用DeepSpeed微调OPT 1.3B模型时遇运算符不匹配错误求助
问题:DeepSpeed微调OPT 1.3B时的CUDA算子编译错误
相关CUDA代码片段
template <typename T> __global__ void moe_res_matmul(T* residual, T* coef, T* mlp_out, int seq_len, int hidden_dim) { constexpr int granularity = 16; constexpr int vals_per_access = granularity / sizeof(T); T* residual_seq = residual + blockIdx.x * hidden_dim; T* mlp_out_seq = mlp_out + blockIdx.x * hidden_dim; for (unsigned tid = threadIdx.x * vals_per_access; tid < hidden_dim; tid += blockDim.x * vals_per_access) { T mlp[vals_per_access]; T res[vals_per_access]; T coef1[vals_per_access]; T coef2[vals_per_access]; mem_access::load_global<granularity>(mlp, mlp_out_seq + tid); mem_access::load_global<granularity>(res, residual_seq + tid); mem_access::load_global<granularity>(coef1, coef + tid); mem_access::load_global<granularity>(coef2, coef + tid + hidden_dim); #pragma unroll for (int idx = 0; idx < vals_per_access; idx++) { mlp[idx] = mlp[idx] * coef2[idx] + res[idx] * coef1[idx]; } mem_access::store_global<granularity>(mlp_out_seq + tid, mlp); } }
错误日志
/.../python3.10/site-packages/deepspeed/ops/csrc/transformer/inference/csrc/gelu.cu(529): error: no operator "*" matches these operands operand types are: __half * __half mlp[idx] = mlp[idx] * coef2[idx] + res[idx] * coef1[idx]; ^ detected during: instantiation of "void moe_res_matmul(T *, T *, T *, int, int) [with T=__half]" at line 547 instantiation of "void launch_moe_res_matmul(T *, T *, T *, int, int, cudaStream_t) [with T=__half]" at line 566
环境依赖
datasets>=2.8.0 sentencepiece>=0.1.97 protobuf==3.20.3 accelerate>=0.15.0 torch>=1.12.0 deepspeed>=0.9.0
解决思路
- 匹配CUDA与PyTorch版本:PyTorch 1.12.0官方推荐搭配CUDA 11.3或11.6,确保系统CUDA版本和PyTorch编译依赖的CUDA版本一致,版本不匹配会导致
__half类型的运算算子无法正确重载。 - 升级DeepSpeed到最新稳定版:0.9.0版本存在该半精度算子的bug,升级到0.12.x及以上版本,官方已修复该类类型运算的兼容性问题。
- 手动修复CUDA算子代码:若无法升级DeepSpeed,找到报错的
gelu.cu文件,将半精度乘法替换为CUDA内置函数:把mlp[idx] * coef2[idx]改为__hmul(mlp[idx], coef2[idx]),res[idx] * coef1[idx]改为__hmul(res[idx], coef1[idx]),__half类型需用专门的CUDA intrinsics函数完成运算。 - 临时禁用半精度训练:调整训练脚本,关闭FP16混合精度(AMP),改用FP32精度训练,绕过该半精度算子的编译问题,但会增加显存占用。
- 从源码重新编译DeepSpeed:卸载当前预编译的DeepSpeed,从源码克隆编译,确保编译时使用的CUDA版本与PyTorch一致,源码编译可解决预编译包的兼容性问题。
内容的提问来源于stack exchange,提问作者coderLMN
相关产品推荐
相关产品推荐

