如何无需导入Torch即可使用已保存的PyTorch模型进行推理?
问题描述
在Slurm集群上执行import torch耗时30-40秒,运维确认是硬件问题,暂无优化方案。对比之下import numpy仅需1秒。
由于需在多个独立任务中重复执行推理,无法缓存Torch导入状态,因此急需无需导入Torch的推理方案——理论上模型推理只是矩阵运算,是否可以用Numpy实现?目标是将总耗时从35秒降至1-2秒。
模型信息
模型通过torch.save(my_network.state_dict(), my_path)保存参数,结构如下:
ActionNN( (conv_1): Conv2d(5, 16, kernel_size=(3, 3), stride=(1, 1)) (conv_2): Conv2d(16, 32, kernel_size=(3, 3), stride=(1, 1)) (conv_3): Conv2d(32, 64, kernel_size=(3, 3), stride=(1, 1)) (norm_layer_1): InstanceNorm2d(16, eps=1e-05, momentum=0.1, affine=False, track_running_stats=False) (norm_layer_2): InstanceNorm2d(32, eps=1e-05, momentum=0.1, affine=False, track_running_stats=False) (norm_layer_3): InstanceNorm2d(64, eps=1e-05, momentum=0.1, affine=False, track_running_stats=False) (gap): AdaptiveAvgPool2d(output_size=(1, 1)) (mlp): Sequential( (0): Linear(in_features=71, out_features=128, bias=True) (1): ReLU() ) (layer): Linear(in_features=128, out_features=210, bias=True) (layer): Linear(in_features=128, out_features=210, bias=True) (layer): Linear(in_features=128, out_features=210, bias=True) (layer): Linear(in_features=128, out_features=210, bias=True) (layer): Linear(in_features=128, out_features=28, bias=True) (layer): Linear(in_features=128, out_features=28, bias=True) (layer): Linear(in_features=128, out_features=28, bias=True) (sigmoid): Sigmoid() (tanh): Tanh() ) Number of parameters: 152284
全连接层单独实现难度低,但该架构包含卷积、InstanceNorm、自适应池化,不确定如何用Numpy实现。
import torch性能分析结果
1226310 function calls (1209639 primitive calls) in 49.994 seconds Ordered by: internal time ncalls tottime percall cumtime percall filename:lineno(function) 1273 21.590 0.017 21.590 0.017 {method 'read' of '_io.BufferedReader' objects} 5276 12.145 0.002 12.145 0.002 {built-in method posix.stat} 1273 7.427 0.006 7.427 0.006 {built-in method io.open_code} 45/25 5.631 0.125 9.939 0.398 {built-in method _imp.create_dynamic} 2 0.564 0.282 0.564 0.282 {built-in method _ctypes.dlopen} 1273 0.288 0.000 0.288 0.000 {built-in method marshal.loads} 17 0.286 0.017 0.286 0.017 {method 'readline' of '_io.BufferedReader' objects} 2809/2753 0.098 0.000 0.546 0.000 {built-in method builtins.__build_class__} 1620/1 0.062 0.000 49.997 49.997 {built-in method builtins.exec} 50145 0.051 0.000 0.119 0.000 {built-in method builtins.getattr} 1159 0.048 0.000 0.115 0.000 inspect.py:3245(signature) 424 0.048 0.000 0.113 0.000 assumptions.py:596(__init__) 13 0.039 0.003 0.039 0.003 {built-in method io.open} 1411 0.035 0.000 0.045 0.000 library.py:71(impl) 1663 0.034 0.000 12.209 0.007 <frozen importlib._bootstrap_external>:1536(find_spec)
解决方案
完全可以用Numpy实现该模型的推理,以下是具体步骤:
1. 预转换模型权重为Numpy格式
这一步只需执行一次,在能快速导入Torch的环境中完成:
import torch import numpy as np # 定义与原模型一致的结构(注意修正同名layer为ModuleList) class ActionNN(torch.nn.Module): def __init__(self): super().__init__() self.conv_1 = torch.nn.Conv2d(5, 16, kernel_size=(3,3), stride=(1,1)) self.conv_2 = torch.nn.Conv2d(16, 32, kernel_size=(3,3), stride=(1,1)) self.conv_3 = torch.nn.Conv2d(32, 64, kernel_size=(3,3), stride=(1,1)) self.norm_layer_1 = torch.nn.InstanceNorm2d(16, eps=1e-05, affine=False) self.norm_layer_2 = torch.nn.InstanceNorm2d(32, eps=1e-05, affine=False) self.norm_layer_3 = torch.nn.InstanceNorm2d(64, eps=1e-05, affine=False) self.gap = torch.nn.AdaptiveAvgPool2d(output_size=(1,1)) self.mlp = torch.nn.Sequential( torch.nn.Linear(71, 128), torch.nn.ReLU() ) self.layers = torch.nn.ModuleList([ torch.nn.Linear(128, 210) for _ in range(4)] + [torch.nn.Linear(128, 28) for _ in range(3)] ) self.sigmoid = torch.nn.Sigmoid() self.tanh = torch.nn.Tanh() # 加载原模型权重并转存为Numpy格式 model = ActionNN() model.load_state_dict(torch.load("your_model_path.pth")) numpy_weights = {} for name, param in model.named_parameters(): numpy_weights[name] = param.cpu().numpy() np.savez("numpy_weights.npz", **numpy_weights)
2. 用Numpy实现各层前向计算
卷积层(Conv2d)
import numpy as np from scipy.signal import convolve2d def conv2d(input_np, weight, bias=None, stride=1): batch_size, in_channels, H, W = input_np.shape out_channels, _, kernel_H, kernel_W = weight.shape out_H = (H - kernel_H) // stride + 1 out_W = (W - kernel_W) // stride + 1 output = np.zeros((batch_size, out_channels, out_H, out_W)) for b in range(batch_size): for out_c in range(out_channels): conv_sum = np.zeros((out_H, out_W)) for in_c in range(in_channels): conv = convolve2d(input_np[b, in_c], weight[out_c, in_c], mode='valid') conv_sum += conv[::stride, ::stride] if bias is not None: conv_sum += bias[out_c] output[b, out_c] = conv_sum return output
InstanceNorm2d
def instance_norm2d(input_np, eps=1e-5): batch_size, channels, H, W = input_np.shape output = np.zeros_like(input_np) for b in range(batch_size): for c in range(channels): channel_data = input_np[b, c] mean = np.mean(channel_data) var = np.var(channel_data) output[b, c] = (channel_data - mean) / np.sqrt(var + eps) return output
AdaptiveAvgPool2d
def adaptive_avg_pool2d(input_np): return np.mean(input_np, axis=(2,3), keepdims=True)
全连接层与激活函数
def linear(input_np, weight, bias=None): output = input_np @ weight.T if bias is not None: output += bias return output def relu(input_np): return np.maximum(input_np, 0) def sigmoid(input_np): return 1 / (1 + np.exp(-input_np)) def tanh(input_np): return np.tanh(input_np)
3. 完整推理流程
# 加载预转换的Numpy权重 weights = np.load("numpy_weights.npz") def predict(input_np): # input_np需符合模型输入格式:(batch, 5, H, W) # 卷积+归一化层 x = conv2d(input_np, weights['conv_1.weight'], weights['conv_1.bias']) x = instance_norm2d(x, eps=1e-5) x = conv2d(x, weights['conv_2.weight'], weights['conv_2.bias']) x = instance_norm2d(x, eps=1e-5) x = conv2d(x, weights['conv_3.weight'], weights['conv_3.bias']) x = instance_norm2d(x, eps=1e-5) # 自适应池化 x = adaptive_avg_pool2d(x) x_flat = x.reshape(x.shape[0], -1) # 注意:原模型MLP输入为71维,需根据实际逻辑补全(比如拼接其他特征) # 示例:若gap输出为64维,需补充7个特征,自行替换为你的实际逻辑 if x_flat.shape[1] != 71: x_flat = np.pad(x_flat, ((0,0),(0,71-x_flat.shape[1])), mode='constant') # MLP层 x = linear(x_flat, weights['mlp.0.weight'], weights['mlp.0.bias']) x = relu(x) # 多输出层 outputs = [] # 前4个210维输出用sigmoid激活 for i in range(4): out = linear(x, weights[f'layers.{i}.weight'], weights[f'layers.{i}.bias']) outputs.append(sigmoid(out)) # 后3个28维输出用tanh激活 for i in range(4,7): out = linear(x, weights[f'layers.{i}.weight'], weights[f'layers.{i}.bias']) outputs.append(tanh(out)) return outputs
额外优化建议
- 若卷积速度慢,可使用
numba对卷积函数做JIT加速,或改用Numpy向量化操作优化 - 也可考虑将模型转为ONNX格式,用
onnxruntime推理——onnxruntime导入速度远快于Torch,且推理效率高 - 若Slurm任务允许,可搭建常驻推理服务,避免重复初始化环境
内容的提问来源于stack exchange,提问作者FluidMechanics Potential Flows
相关产品推荐
相关产品推荐

