如何提升TensorFlow图像目标识别推理脚本的性能?
问题描述
我编写了一个使用TensorFlow识别JPEG图像中目标的脚本,发现推理阶段耗时极长。分析10张单张小于2MB的JPEG图像耗时达222.91秒,相关代码如下:
N_CHANNELS = 3 def load_image_into_numpy_array(image): """ Converts a PIL image into a numpy array (height x width x channels). :param image: PIL image :return: numpy array """ (width, height) = image.size return np.array(image.getdata()) \ .reshape((height, width, N_CHANNELS)).astype(np.uint8) def process_output(classes, scores, boxes, category_index): """ Processes classes, scores, and boxes, gathering in a list of ObjectResult. :param classes: list of class id :param scores: list of scores :param boxes: list of boxes :param category_index: label dictionary :return: list of ObjectResult """ results = [] for clazz, score, box in zip(classes, scores, boxes): if score > 0.0: label = category_index[clazz][NAME_KEY] obj_result = ObjectResult(label, score, box) results.append(obj_result) return results # Inference tic = time.perf_counter() IMAGE_NP_KEY = 'image_np' RESULTS_KEY = 'results' file_result_dict = {} for filename in TEST_IMAGES: image_np = load_image_into_numpy_array(Image.open(filename)) output_dict = run_inference(graph, image_np) results = process_output(output_dict[DETECTION_CLASSES_KEY], output_dict[DETECTION_SCORES_KEY], output_dict[DETECTION_BOXES_KEY], category_index) file_result_dict[filename] = { IMAGE_NP_KEY: image_np, RESULTS_KEY: results } toc = time.perf_counter() print("Inference completed in", round(toc - tic, 2), "seconds")
环境信息:
- TensorFlow版本:1.14.0
- CPU:3.6 GHz 10核Intel Core i9
- AMD GPU(Radeon Pro 5700 8 GB)不被TensorFlow支持
- 内存:16GB 2667 MHz DDR4
- 操作系统:macOS 12.4
- 在Conda环境中运行
请问如何优化该脚本的性能?
优化方案
1. 升级TensorFlow版本
直接升级到TensorFlow 2.x系列稳定版本(如2.15.x),旧版1.14.0在CPU推理、macOS适配等方面存在性能短板,新版本针对Intel CPU向量指令、计算图执行效率有明显优化。
2. 优化图像加载逻辑
替换低效的load_image_into_numpy_array方法,利用PIL直接转numpy数组的原生实现,避免逐像素遍历的开销:
def load_image_into_numpy_array(image): return np.array(image).astype(np.uint8)
也可改用TensorFlow内置的tf.io.read_file+tf.image.decode_jpeg加载图像,借助TF的并行预处理能力进一步提速。
3. 充分利用CPU多核资源
手动配置TensorFlow线程数,最大化10核CPU的利用率:
# TensorFlow 1.x配置 config = tf.ConfigProto( intra_op_parallelism_threads=10, # 设为CPU核心数 inter_op_parallelism_threads=2 ) sess = tf.Session(graph=graph, config=config) # TensorFlow 2.x配置 tf.config.threading.set_intra_op_parallelism_threads(10) tf.config.threading.set_inter_op_parallelism_threads(2)
4. 模型量化与图优化
- INT8量化:用
tf.lite.TFLiteConverter将模型转为INT8量化的TFLite模型,CPU推理速度可提升2-4倍。 - 冻结优化计算图:通过
tf.graph_util.convert_variables_to_constants冻结图,再移除训练节点,减少冗余计算。
5. 改为批量推理
将单张循环推理改为批量输入,避免重复的图初始化和数据传输开销:
# 批量加载并统一尺寸(确保图像尺寸与模型输入一致) INPUT_WIDTH, INPUT_HEIGHT = 640, 640 # 替换为你的模型输入尺寸 batch_images = [np.array(Image.open(f).resize((INPUT_WIDTH, INPUT_HEIGHT))).astype(np.uint8) for f in TEST_IMAGES] batch_images_np = np.stack(batch_images) # 单次批量推理 output_dict = run_inference(graph, batch_images_np) # 批量解析结果 for idx, filename in enumerate(TEST_IMAGES): results = process_output( output_dict[DETECTION_CLASSES_KEY][idx], output_dict[DETECTION_SCORES_KEY][idx], output_dict[DETECTION_BOXES_KEY][idx], category_index ) file_result_dict[filename] = { RESULTS_KEY: results } # 若无需原图可移除IMAGE_NP_KEY
6. 清理冗余存储
移除file_result_dict中不必要的image_np存储,减少内存占用,避免内存交换拖慢性能。
7. 启用macOS Metal加速
安装tensorflow-metal插件(仅支持TF 2.5+),利用Apple Metal框架提升推理性能,即使AMD GPU也能获得加速:
pip install tensorflow-metal
内容的提问来源于stack exchange,提问作者Raptor
相关产品推荐
相关产品推荐

