TorchScript部署ALBEF模型遇设备不匹配RuntimeError求助
问题:ALBEF模型TorchScript导出后C++/Python推理均报设备不匹配错误
问题背景
在Python中训练了基于ALBEF的模型,为提升推理效率计划在C部署。使用torch.jit.trace导出为.pt文件后,无论是C加载推理,还是Python重新加载该模型推理,都触发RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu。
C++推理代码
if (torch::cuda::is_available()) { n_model = torch::jit::load("/home/lzh/Storage4/lzh/deepmodel/model_scripted.pt",torch::kCUDA); std::cout << torch::cuda::device_count() << std::endl; } else { std::cerr << "No CUDA devices available, cannot move model to GPU." << std::endl; } torch::Tensor inputs = torch::from_blob(fre, {1, 4,300, 201}, torch::kFloat).to(torch::kCUDA); std::cout << inputs.device() << std::endl; textInput.input_ids.to(torch::kCUDA); textInput.attention_mask.to(torch::kCUDA); torch::Tensor out_tensor = n_model.forward({inputs,textInput.input_ids,textInput.attention_mask}).toTensor();
报错信息
The following operation failed in the TorchScript interpreter. Traceback of TorchScript, serialized code (most recent call last): File "code/__torch__/models/model_somatic.py", line 14, in forward cls_head = self.cls_head ALBEF = self.ALBEF _0 = (ALBEF).forward(image, input_ids, attention_mask, ) ~~~~~~~~~~~~~~ <--- HERE return (cls_head).forward(_0, ) class ALBEF(Module): File "code/__torch__/models/model_somatic.py", line 35, in forward _5 = torch.ones([_3, int(_4)], dtype=4, layout=None, device=torch.device("cpu"), pin_memory=False) encoder_attention_mask = torch.to(_5, dtype=4, layout=0, device=torch.device("cpu")) _6 = (text_encoder).forward(input_ids, attention_mask, _1, encoder_attention_mask, ) ~~~~~~~~~~~~~~~~~~~~~ <--- HERE _7 = torch.slice(_6, 0, 0, 9223372036854775807) input = torch.slice(torch.select(_7, 1, 0), 1, 0, 9223372036854775807) File "code/__torch__/models/xbert.py", line 19, in forward cls = self.cls bert0 = self.bert _0 = (bert0).forward(input_ids, attention_mask, argument_3, encoder_attention_mask, ) ~~~~~~~~~~~~~~ <--- HERE _1 = (cls).forward(weight, _0, ) return _0 File "code/__torch__/models/xbert.py", line 50, in forward _8 = torch.to(encoder_extended_attention_mask, 6) attention_mask1 = torch.mul(torch.rsub(_8, 1.), CONSTANTS.c3) _9 = (embeddings).forward(input_ids, input, ) ~~~~~~~~~~~~~~~~~~~ <--- HERE _10 = (encoder).forward(_9, attention_mask0, argument_3, attention_mask1, ) return _10 File "code/__torch__/models/xbert.py", line 78, in forward input0 = torch.slice(_12, 1, 0, _11) _13 = (word_embeddings).forward(input_ids, ) _14 = (token_type_embeddings).forward(input, ) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ <--- HERE embeddings = torch.add(_13, _14) _15 = (position_embeddings).forward(input0, ) File "code/__torch__/torch/nn/modules/sparse/___torch_mangle_164.py", line 10, in forward input: Tensor) -> Tensor: weight = self.weight return torch.embedding(weight, input) ~~~~~~~~~~~~~~~ <--- HERE Traceback of TorchScript, original code (most recent call last): /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/nn/functional.py(2044): embedding /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/nn/modules/sparse.py(158): forward /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/nn/modules/module.py(1090): _slow_forward /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/nn/modules/module.py(1102): _call_impl /home/lzh/ALBEF/models/xbert.py(207): forward /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/nn/modules/module.py(1090): _slow_forward /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/nn/modules/module.py(1102): _call_impl /home/lzh/ALBEF/models/xbert.py(1046): forward /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/nn/modules/module.py(1090): _slow_forward /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/nn/modules/module.py(1102): _call_impl /home/lzh/ALBEF/models/xbert.py(1400): forward /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/nn/modules/module.py(1090): _slow_forward /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/nn/modules/module.py(1102): _call_impl /home/lzh/ALBEF/models/model_somatic.py(47): forward /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/nn/modules/module.py(1090): _slow_forward /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/nn/modules/module.py(1102): _call_impl /home/lzh/ALBEF/models/model_somatic.py(90): forward /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/nn/modules/module.py(1090): _slow_forward /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/nn/modules/module.py(1102): _call_impl /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/jit/_trace.py(958): trace_module /home/lzh/miniconda3/envs/albef/lib/python3.8/site-packages/torch/jit/_trace.py(741): trace /home/lzh/ALBEF/checkpoint.py(46): main /home/lzh/ALBEF/checkpoint.py(76): <module> RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu! (when checking argument for argument index in method wrapper_CUDA__index_select)
Python测试代码(加载导出模型也报错)
image = torch.rand(16,4,300,201) text1 = torch.rand(16,25).long() text2 = torch.rand(16, 25).long() traced_script_module = torch.jit.trace(model, (image,text1,text2)) traced_script_module.save('model_scripted.pt') device=torch.device("cuda:0") text = torch.ones((1,25)) text = text.long().to(device) image = torch.ones((1,4,300,201)).to(device) model = torch.jit.load('model_scripted.pt', map_location=torch.device('cuda')) model.eval() for param in model.parameters(): if param.device.type == 'cuda': print('cuda') print(image.device) print(text.device) out = model(image,text,text)
已尝试方案
尝试过常规设备匹配调试方案,但未解决问题。
解决方法
从报错栈可见,模型forward过程中硬编码了device=torch.device("cpu")创建张量(如torch.ones([_3, int(_4)], ..., device=torch.device("cpu"))),导致该张量始终在CPU,与其他CUDA设备的张量运算时冲突。
解决步骤:
修改模型代码:找到
model_somatic.py中创建encoder_attention_mask的代码,去掉硬编码的device="cpu",改为根据输入张量的设备自动设置:# 原问题代码 _5 = torch.ones([_3, int(_4)], device=torch.device("cpu")) # 修改后 _5 = torch.ones([_3, int(_4)], device=input_ids.device)确保所有动态创建的张量都与输入/模型参数同设备。
重新导出模型:在Python中将模型移到CUDA设备后再执行trace,避免trace时记录CPU设备的默认行为:
device = torch.device("cuda:0") model.to(device) model.eval() # 用CUDA输入做trace image = torch.rand(16,4,300,201).to(device) text1 = torch.rand(16,25).long().to(device) text2 = torch.rand(16,25).long().to(device) traced_script_module = torch.jit.trace(model, (image, text1, text2)) traced_script_module.save('model_scripted_cuda.pt')C++端修正张量赋值:原代码中
textInput.input_ids.to(torch::kCUDA)仅返回新张量,未修改原变量,需改为:textInput.input_ids = textInput.input_ids.to(torch::kCUDA); textInput.attention_mask = textInput.attention_mask.to(torch::kCUDA);
完成以上步骤后,重新加载模型推理即可解决设备不匹配问题。
内容的提问来源于stack exchange,提问作者LZH
相关产品推荐
相关产品推荐

