You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在不批量采样的情况下加速深度学习推理?

逐样本顺序推理提速方案与最佳实践

问题背景

出于算法需求,必须采用逐样本顺序执行推理,而非批量采样。使用的DNN结构如下:

Model: "sequential"
_________________________________________________________________
 Layer (type)                Output Shape              Param #   
=================================================================
 dense (Dense)               (None, 64)                2624      
                                                                 
 dense_1 (Dense)             (None, 128)               8320      
                                                                 
 dense_2 (Dense)             (None, 128)               16512     
                                                                 
 dense_3 (Dense)             (None, 1)                 129       
                                                                 
=================================================================
Total params: 27585 (107.75 KB)
Trainable params: 27585 (107.75 KB)
Non-trainable params: 0 (0.00 Byte)
_________________________________________________________________

批量推理测试(耗时1.34秒):

start = time.time()
y_pred = model.predict(x_test_feature_space)
end = time.time()
print("Exeuction time: ", end - start)
print("Execution time per sample: ", (end-start)/len(x_test_feature_space))

# 输出:
# 672/672 [==============================] - 1s 1ms/step
# Exeuction time:  1.345379114151001
# Execution time per sample:  6.26194607470794e-05

逐样本推理测试(耗时24分钟):

y_pred = []
timings = []

start_total = time.time()
for x in x_test_feature_space:
    start = time.time()
    y_pred.append(model.predict(np.array([x,]), verbose=False))
    end = time.time()
    timings.append(end-start)
end_total = time.time()

print("Exeuction time: ", end_total - start_total)
print("Average time per sample: ", sum(timings)/len(timings))

# 输出:
# Exeuction time:  1467.2437105178833
# Average time per sample:  0.06825089837496631

从耗时差异能看出,推理本身的计算成本远低于每次调用model.predict()带来的额外开销。现需解决:如何提升逐样本顺序推理速度?相关最佳实践有哪些?

可行提速方法与最佳实践

1. 替换model.predict()为底层推理接口

model.predict()每次调用都会触发输入校验、张量转换、会话初始化等冗余操作。直接使用模型的__call__方法(即model(x)),跳过这些包装逻辑:

# 替换循环内的predict调用
y_pred.append(model(np.array([x,])))

该方式直接执行前向传播,能大幅降低单样本推理的额外开销。

2. 预转换输入格式,减少循环内冗余操作

逐样本循环中反复将x转为np.array([x,])会产生重复的内存分配与格式转换成本。可提前将整个测试集转为模型期望的张量格式,循环中直接取单样本:

# 提前转换数据集为张量(以TensorFlow为例)
import tensorflow as tf
x_tensor = tf.convert_to_tensor(x_test_feature_space)

y_pred = []
start_total = time.time()
for i in range(len(x_test_feature_space)):
    # 直接取单样本切片
    y_pred.append(model(x_tensor[i:i+1]))
end_total = time.time()

3. 启用推理优化模式

开启模型的推理专用模式,关闭训练相关操作(如Dropout、BatchNorm的训练态),同时用计算图优化前向路径:

# TensorFlow中启用推理模式
model.trainable = False

# 用tf.function装饰推理函数,转为计算图执行
@tf.function
def predict_single(x):
    return model(x)

# 循环内调用优化后的函数
y_pred.append(predict_single(np.array([x,])))

tf.function会将Python代码转为TensorFlow计算图,减少解释执行的开销,多次调用时性能提升明显。

4. 模型轻量化与量化

  • 精度量化:将模型从FP32转为FP16或INT8,减少计算量与内存带宽占用,同时不显著降低精度。TensorFlow可通过tf.lite.TFLiteConverter完成量化转换,用TFLite解释器执行推理。
  • 结构轻量化:如果允许,用更小的模型蒸馏原模型的知识,减少单样本推理的计算量(需额外训练步骤)。

5. 简化循环内操作

将循环内的非推理操作(如计时、列表扩容)尽量移出或简化:

# 预分配结果列表空间,避免频繁扩容
y_pred = [None] * len(x_test_feature_space)
for i in range(len(x_test_feature_space)):
    y_pred[i] = model(x_tensor[i:i+1])

6. 使用专用推理引擎

将模型导出为ONNX格式,使用ONNX Runtime、TensorRT等专用推理引擎执行单样本推理。这类引擎针对推理场景做了深度优化,能有效降低单样本延迟:

# 示例:使用ONNX Runtime
import onnxruntime as ort
import tf2onnx

# 将Keras模型转为ONNX格式
onnx_model, _ = tf2onnx.convert.from_keras(model, opset=13)
with open("model.onnx", "wb") as f:
    f.write(onnx_model.SerializeToString())

# 加载模型并逐样本推理
sess = ort.InferenceSession("model.onnx")
input_name = sess.get_inputs()[0].name

y_pred = []
for x in x_test_feature_space:
    pred = sess.run(None, {input_name: np.array([x,])})
    y_pred.append(pred)

内容的提问来源于stack exchange,提问作者Polo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 13:16:26