LibTorch模型无法常驻CUDA设备问题求助
LibTorch模型无法常驻CUDA设备的问题
无法将模型放置并常驻在CUDA设备上。当传入已在CUDA上的张量时,会触发**"found at least two devices, cpu and cuda"**错误。
我是否遗漏了在LibTorch中将模型部署到CUDA设备的简单方法?找不到解决方案。
核心问题代码示例
当张量已在CUDA上时,传入模型会报错:
auto the_tensor = torch::rand({42, 427}).to(device); std::cout << net.forward(the_tensor).to(device); terminate called after throwing an instance of 'c10::Error' what(): Expected all tensors to be on the same device, but found at least two devices, cpu and cuda:0! (when checking argument for argument mat1 in method wrapper_addmm)
若张量不在CUDA上,可正常在CUDA运行模型:
auto the_tensor = torch::rand({42, 427}); std::cout << net.forward(the_tensor).to(device);
将张量转回CPU也不会报错,但我的脚本中有大量已在CUDA上的张量,不想来回转移,因此寻求将模型永久部署在CUDA设备的方法。
尝试过的无效方法
net.to(device); net->to(device); Critic_Net().to(device);
只有在forward调用后添加.to(device)才能正常运行,以下是完整可复现代码:
#include <torch/torch.h> using namespace torch::indexing; torch::Device device(torch::kCUDA); struct Critic_Net : torch::nn::Module { torch::Tensor next_state_batch__sampled_action; public: Critic_Net() { lin1 = torch::nn::Linear(427, 42); lin2 = torch::nn::Linear(42, 286); lin3 = torch::nn::Linear(286, 1); } torch::Tensor forward(torch::Tensor next_state_batch__sampled_action) { auto h = next_state_batch__sampled_action; h = torch::relu(lin1->forward(h)); h = torch::tanh(lin2->forward(h)); h = lin3->forward(h); return torch::nan_to_num(h); } torch::nn::Linear lin1{nullptr}, lin2{nullptr}, lin3{nullptr}; }; auto net = Critic_Net(); int main() { net.to(device); auto the_tensor = torch::rand({42, 427}).to(device); std::cout << net.forward(the_tensor).to(device); }
内容的提问来源于stack exchange,提问作者Ant
相关产品推荐
相关产品推荐

