You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何无需导入Torch即可使用已保存的PyTorch模型进行推理?

问题描述

在Slurm集群上执行import torch耗时30-40秒,运维确认是硬件问题,暂无优化方案。对比之下import numpy仅需1秒。

由于需在多个独立任务中重复执行推理,无法缓存Torch导入状态,因此急需无需导入Torch的推理方案——理论上模型推理只是矩阵运算,是否可以用Numpy实现?目标是将总耗时从35秒降至1-2秒。

模型信息

模型通过torch.save(my_network.state_dict(), my_path)保存参数,结构如下:

ActionNN(
  (conv_1): Conv2d(5, 16, kernel_size=(3, 3), stride=(1, 1))
  (conv_2): Conv2d(16, 32, kernel_size=(3, 3), stride=(1, 1))
  (conv_3): Conv2d(32, 64, kernel_size=(3, 3), stride=(1, 1))
  (norm_layer_1): InstanceNorm2d(16, eps=1e-05, momentum=0.1, affine=False, track_running_stats=False)
  (norm_layer_2): InstanceNorm2d(32, eps=1e-05, momentum=0.1, affine=False, track_running_stats=False)
  (norm_layer_3): InstanceNorm2d(64, eps=1e-05, momentum=0.1, affine=False, track_running_stats=False)
  (gap): AdaptiveAvgPool2d(output_size=(1, 1))
  (mlp): Sequential(
    (0): Linear(in_features=71, out_features=128, bias=True)
    (1): ReLU()
  )
  (layer): Linear(in_features=128, out_features=210, bias=True)
  (layer): Linear(in_features=128, out_features=210, bias=True)
  (layer): Linear(in_features=128, out_features=210, bias=True)
  (layer): Linear(in_features=128, out_features=210, bias=True)
  (layer): Linear(in_features=128, out_features=28, bias=True)
  (layer): Linear(in_features=128, out_features=28, bias=True)
  (layer): Linear(in_features=128, out_features=28, bias=True)
  (sigmoid): Sigmoid()
  (tanh): Tanh()
)
Number of parameters: 152284

全连接层单独实现难度低,但该架构包含卷积、InstanceNorm、自适应池化,不确定如何用Numpy实现。

import torch性能分析结果

1226310 function calls (1209639 primitive calls) in 49.994 seconds

   Ordered by: internal time

   ncalls  tottime  percall  cumtime  percall filename:lineno(function)
     1273   21.590    0.017   21.590    0.017 {method 'read' of '_io.BufferedReader' objects}
     5276   12.145    0.002   12.145    0.002 {built-in method posix.stat}
     1273    7.427    0.006    7.427    0.006 {built-in method io.open_code}
    45/25    5.631    0.125    9.939    0.398 {built-in method _imp.create_dynamic}
        2    0.564    0.282    0.564    0.282 {built-in method _ctypes.dlopen}
     1273    0.288    0.000    0.288    0.000 {built-in method marshal.loads}
       17    0.286    0.017    0.286    0.017 {method 'readline' of '_io.BufferedReader' objects}
2809/2753    0.098    0.000    0.546    0.000 {built-in method builtins.__build_class__}
   1620/1    0.062    0.000   49.997   49.997 {built-in method builtins.exec}
    50145    0.051    0.000    0.119    0.000 {built-in method builtins.getattr}
     1159    0.048    0.000    0.115    0.000 inspect.py:3245(signature)
      424    0.048    0.000    0.113    0.000 assumptions.py:596(__init__)
       13    0.039    0.003    0.039    0.003 {built-in method io.open}
     1411    0.035    0.000    0.045    0.000 library.py:71(impl)
     1663    0.034    0.000   12.209    0.007 <frozen importlib._bootstrap_external>:1536(find_spec)

解决方案

完全可以用Numpy实现该模型的推理,以下是具体步骤:

1. 预转换模型权重为Numpy格式

这一步只需执行一次,在能快速导入Torch的环境中完成:

import torch
import numpy as np

# 定义与原模型一致的结构(注意修正同名layer为ModuleList)
class ActionNN(torch.nn.Module):
    def __init__(self):
        super().__init__()
        self.conv_1 = torch.nn.Conv2d(5, 16, kernel_size=(3,3), stride=(1,1))
        self.conv_2 = torch.nn.Conv2d(16, 32, kernel_size=(3,3), stride=(1,1))
        self.conv_3 = torch.nn.Conv2d(32, 64, kernel_size=(3,3), stride=(1,1))
        self.norm_layer_1 = torch.nn.InstanceNorm2d(16, eps=1e-05, affine=False)
        self.norm_layer_2 = torch.nn.InstanceNorm2d(32, eps=1e-05, affine=False)
        self.norm_layer_3 = torch.nn.InstanceNorm2d(64, eps=1e-05, affine=False)
        self.gap = torch.nn.AdaptiveAvgPool2d(output_size=(1,1))
        self.mlp = torch.nn.Sequential(
            torch.nn.Linear(71, 128),
            torch.nn.ReLU()
        )
        self.layers = torch.nn.ModuleList([
            torch.nn.Linear(128, 210) for _ in range(4)] +
            [torch.nn.Linear(128, 28) for _ in range(3)]
        )
        self.sigmoid = torch.nn.Sigmoid()
        self.tanh = torch.nn.Tanh()

# 加载原模型权重并转存为Numpy格式
model = ActionNN()
model.load_state_dict(torch.load("your_model_path.pth"))

numpy_weights = {}
for name, param in model.named_parameters():
    numpy_weights[name] = param.cpu().numpy()

np.savez("numpy_weights.npz", **numpy_weights)

2. 用Numpy实现各层前向计算

卷积层(Conv2d)

import numpy as np
from scipy.signal import convolve2d

def conv2d(input_np, weight, bias=None, stride=1):
    batch_size, in_channels, H, W = input_np.shape
    out_channels, _, kernel_H, kernel_W = weight.shape
    
    out_H = (H - kernel_H) // stride + 1
    out_W = (W - kernel_W) // stride + 1
    output = np.zeros((batch_size, out_channels, out_H, out_W))
    
    for b in range(batch_size):
        for out_c in range(out_channels):
            conv_sum = np.zeros((out_H, out_W))
            for in_c in range(in_channels):
                conv = convolve2d(input_np[b, in_c], weight[out_c, in_c], mode='valid')
                conv_sum += conv[::stride, ::stride]
            if bias is not None:
                conv_sum += bias[out_c]
            output[b, out_c] = conv_sum
    return output

InstanceNorm2d

def instance_norm2d(input_np, eps=1e-5):
    batch_size, channels, H, W = input_np.shape
    output = np.zeros_like(input_np)
    
    for b in range(batch_size):
        for c in range(channels):
            channel_data = input_np[b, c]
            mean = np.mean(channel_data)
            var = np.var(channel_data)
            output[b, c] = (channel_data - mean) / np.sqrt(var + eps)
    return output

AdaptiveAvgPool2d

def adaptive_avg_pool2d(input_np):
    return np.mean(input_np, axis=(2,3), keepdims=True)

全连接层与激活函数

def linear(input_np, weight, bias=None):
    output = input_np @ weight.T
    if bias is not None:
        output += bias
    return output

def relu(input_np):
    return np.maximum(input_np, 0)

def sigmoid(input_np):
    return 1 / (1 + np.exp(-input_np))

def tanh(input_np):
    return np.tanh(input_np)

3. 完整推理流程

# 加载预转换的Numpy权重
weights = np.load("numpy_weights.npz")

def predict(input_np):
    # input_np需符合模型输入格式:(batch, 5, H, W)
    # 卷积+归一化层
    x = conv2d(input_np, weights['conv_1.weight'], weights['conv_1.bias'])
    x = instance_norm2d(x, eps=1e-5)
    
    x = conv2d(x, weights['conv_2.weight'], weights['conv_2.bias'])
    x = instance_norm2d(x, eps=1e-5)
    
    x = conv2d(x, weights['conv_3.weight'], weights['conv_3.bias'])
    x = instance_norm2d(x, eps=1e-5)
    
    # 自适应池化
    x = adaptive_avg_pool2d(x)
    x_flat = x.reshape(x.shape[0], -1)
    
    # 注意:原模型MLP输入为71维,需根据实际逻辑补全(比如拼接其他特征)
    # 示例:若gap输出为64维,需补充7个特征,自行替换为你的实际逻辑
    if x_flat.shape[1] != 71:
        x_flat = np.pad(x_flat, ((0,0),(0,71-x_flat.shape[1])), mode='constant')
    
    # MLP层
    x = linear(x_flat, weights['mlp.0.weight'], weights['mlp.0.bias'])
    x = relu(x)
    
    # 多输出层
    outputs = []
    # 前4个210维输出用sigmoid激活
    for i in range(4):
        out = linear(x, weights[f'layers.{i}.weight'], weights[f'layers.{i}.bias'])
        outputs.append(sigmoid(out))
    # 后3个28维输出用tanh激活
    for i in range(4,7):
        out = linear(x, weights[f'layers.{i}.weight'], weights[f'layers.{i}.bias'])
        outputs.append(tanh(out))
    
    return outputs

额外优化建议

  • 若卷积速度慢,可使用numba对卷积函数做JIT加速,或改用Numpy向量化操作优化
  • 也可考虑将模型转为ONNX格式,用onnxruntime推理——onnxruntime导入速度远快于Torch,且推理效率高
  • 若Slurm任务允许,可搭建常驻推理服务,避免重复初始化环境

内容的提问来源于stack exchange,提问作者FluidMechanics Potential Flows

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 15:42:02