如何在不批量采样的情况下加速深度学习推理?
逐样本顺序推理提速方案与最佳实践
问题背景
出于算法需求,必须采用逐样本顺序执行推理,而非批量采样。使用的DNN结构如下:
Model: "sequential" _________________________________________________________________ Layer (type) Output Shape Param # ================================================================= dense (Dense) (None, 64) 2624 dense_1 (Dense) (None, 128) 8320 dense_2 (Dense) (None, 128) 16512 dense_3 (Dense) (None, 1) 129 ================================================================= Total params: 27585 (107.75 KB) Trainable params: 27585 (107.75 KB) Non-trainable params: 0 (0.00 Byte) _________________________________________________________________
批量推理测试(耗时1.34秒):
start = time.time() y_pred = model.predict(x_test_feature_space) end = time.time() print("Exeuction time: ", end - start) print("Execution time per sample: ", (end-start)/len(x_test_feature_space)) # 输出: # 672/672 [==============================] - 1s 1ms/step # Exeuction time: 1.345379114151001 # Execution time per sample: 6.26194607470794e-05
逐样本推理测试(耗时24分钟):
y_pred = [] timings = [] start_total = time.time() for x in x_test_feature_space: start = time.time() y_pred.append(model.predict(np.array([x,]), verbose=False)) end = time.time() timings.append(end-start) end_total = time.time() print("Exeuction time: ", end_total - start_total) print("Average time per sample: ", sum(timings)/len(timings)) # 输出: # Exeuction time: 1467.2437105178833 # Average time per sample: 0.06825089837496631
从耗时差异能看出,推理本身的计算成本远低于每次调用model.predict()带来的额外开销。现需解决:如何提升逐样本顺序推理速度?相关最佳实践有哪些?
可行提速方法与最佳实践
1. 替换model.predict()为底层推理接口
model.predict()每次调用都会触发输入校验、张量转换、会话初始化等冗余操作。直接使用模型的__call__方法(即model(x)),跳过这些包装逻辑:
# 替换循环内的predict调用 y_pred.append(model(np.array([x,])))
该方式直接执行前向传播,能大幅降低单样本推理的额外开销。
2. 预转换输入格式,减少循环内冗余操作
逐样本循环中反复将x转为np.array([x,])会产生重复的内存分配与格式转换成本。可提前将整个测试集转为模型期望的张量格式,循环中直接取单样本:
# 提前转换数据集为张量(以TensorFlow为例) import tensorflow as tf x_tensor = tf.convert_to_tensor(x_test_feature_space) y_pred = [] start_total = time.time() for i in range(len(x_test_feature_space)): # 直接取单样本切片 y_pred.append(model(x_tensor[i:i+1])) end_total = time.time()
3. 启用推理优化模式
开启模型的推理专用模式,关闭训练相关操作(如Dropout、BatchNorm的训练态),同时用计算图优化前向路径:
# TensorFlow中启用推理模式 model.trainable = False # 用tf.function装饰推理函数,转为计算图执行 @tf.function def predict_single(x): return model(x) # 循环内调用优化后的函数 y_pred.append(predict_single(np.array([x,])))
tf.function会将Python代码转为TensorFlow计算图,减少解释执行的开销,多次调用时性能提升明显。
4. 模型轻量化与量化
- 精度量化:将模型从FP32转为FP16或INT8,减少计算量与内存带宽占用,同时不显著降低精度。TensorFlow可通过
tf.lite.TFLiteConverter完成量化转换,用TFLite解释器执行推理。 - 结构轻量化:如果允许,用更小的模型蒸馏原模型的知识,减少单样本推理的计算量(需额外训练步骤)。
5. 简化循环内操作
将循环内的非推理操作(如计时、列表扩容)尽量移出或简化:
# 预分配结果列表空间,避免频繁扩容 y_pred = [None] * len(x_test_feature_space) for i in range(len(x_test_feature_space)): y_pred[i] = model(x_tensor[i:i+1])
6. 使用专用推理引擎
将模型导出为ONNX格式,使用ONNX Runtime、TensorRT等专用推理引擎执行单样本推理。这类引擎针对推理场景做了深度优化,能有效降低单样本延迟:
# 示例:使用ONNX Runtime import onnxruntime as ort import tf2onnx # 将Keras模型转为ONNX格式 onnx_model, _ = tf2onnx.convert.from_keras(model, opset=13) with open("model.onnx", "wb") as f: f.write(onnx_model.SerializeToString()) # 加载模型并逐样本推理 sess = ort.InferenceSession("model.onnx") input_name = sess.get_inputs()[0].name y_pred = [] for x in x_test_feature_space: pred = sess.run(None, {input_name: np.array([x,])}) y_pred.append(pred)
内容的提问来源于stack exchange,提问作者Polo
相关产品推荐
相关产品推荐

