You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升TensorFlow图像目标识别推理脚本的性能?

问题描述

我编写了一个使用TensorFlow识别JPEG图像中目标的脚本,发现推理阶段耗时极长。分析10张单张小于2MB的JPEG图像耗时达222.91秒,相关代码如下:

N_CHANNELS = 3
def load_image_into_numpy_array(image):
    """
    Converts a PIL image into a numpy array (height x width x channels).
    :param image: PIL image
    :return: numpy array
    """
    (width, height) = image.size
    return np.array(image.getdata()) \
        .reshape((height, width, N_CHANNELS)).astype(np.uint8)

def process_output(classes, scores, boxes, category_index):
    """
    Processes classes, scores, and boxes, gathering in a list of ObjectResult.
    :param classes: list of class id
    :param scores: list of scores
    :param boxes: list of boxes
    :param category_index: label dictionary
    :return: list of ObjectResult
    """
    results = []

    for clazz, score, box in zip(classes, scores, boxes):
        if score > 0.0:
            label = category_index[clazz][NAME_KEY]
            obj_result = ObjectResult(label, score, box)
            results.append(obj_result)
    
    return results

# Inference
tic = time.perf_counter()
IMAGE_NP_KEY = 'image_np'
RESULTS_KEY = 'results'

file_result_dict = {}

for filename in TEST_IMAGES:
    image_np = load_image_into_numpy_array(Image.open(filename))
    
    output_dict = run_inference(graph, image_np)
 
    results = process_output(output_dict[DETECTION_CLASSES_KEY],
                             output_dict[DETECTION_SCORES_KEY],
                             output_dict[DETECTION_BOXES_KEY],
                             category_index)

    file_result_dict[filename] = { IMAGE_NP_KEY: image_np, RESULTS_KEY: results }
toc = time.perf_counter()
print("Inference completed in", round(toc - tic, 2), "seconds")

环境信息:

  • TensorFlow版本:1.14.0
  • CPU:3.6 GHz 10核Intel Core i9
  • AMD GPU(Radeon Pro 5700 8 GB)不被TensorFlow支持
  • 内存:16GB 2667 MHz DDR4
  • 操作系统:macOS 12.4
  • 在Conda环境中运行

请问如何优化该脚本的性能?


优化方案

1. 升级TensorFlow版本

直接升级到TensorFlow 2.x系列稳定版本(如2.15.x),旧版1.14.0在CPU推理、macOS适配等方面存在性能短板,新版本针对Intel CPU向量指令、计算图执行效率有明显优化。

2. 优化图像加载逻辑

替换低效的load_image_into_numpy_array方法,利用PIL直接转numpy数组的原生实现,避免逐像素遍历的开销:

def load_image_into_numpy_array(image):
    return np.array(image).astype(np.uint8)

也可改用TensorFlow内置的tf.io.read_file+tf.image.decode_jpeg加载图像,借助TF的并行预处理能力进一步提速。

3. 充分利用CPU多核资源

手动配置TensorFlow线程数,最大化10核CPU的利用率:

# TensorFlow 1.x配置
config = tf.ConfigProto(
    intra_op_parallelism_threads=10,  # 设为CPU核心数
    inter_op_parallelism_threads=2
)
sess = tf.Session(graph=graph, config=config)

# TensorFlow 2.x配置
tf.config.threading.set_intra_op_parallelism_threads(10)
tf.config.threading.set_inter_op_parallelism_threads(2)

4. 模型量化与图优化

  • INT8量化:用tf.lite.TFLiteConverter将模型转为INT8量化的TFLite模型,CPU推理速度可提升2-4倍。
  • 冻结优化计算图:通过tf.graph_util.convert_variables_to_constants冻结图,再移除训练节点,减少冗余计算。

5. 改为批量推理

将单张循环推理改为批量输入,避免重复的图初始化和数据传输开销:

# 批量加载并统一尺寸(确保图像尺寸与模型输入一致)
INPUT_WIDTH, INPUT_HEIGHT = 640, 640  # 替换为你的模型输入尺寸
batch_images = [np.array(Image.open(f).resize((INPUT_WIDTH, INPUT_HEIGHT))).astype(np.uint8) for f in TEST_IMAGES]
batch_images_np = np.stack(batch_images)

# 单次批量推理
output_dict = run_inference(graph, batch_images_np)

# 批量解析结果
for idx, filename in enumerate(TEST_IMAGES):
    results = process_output(
        output_dict[DETECTION_CLASSES_KEY][idx],
        output_dict[DETECTION_SCORES_KEY][idx],
        output_dict[DETECTION_BOXES_KEY][idx],
        category_index
    )
    file_result_dict[filename] = { RESULTS_KEY: results }  # 若无需原图可移除IMAGE_NP_KEY

6. 清理冗余存储

移除file_result_dict中不必要的image_np存储,减少内存占用,避免内存交换拖慢性能。

7. 启用macOS Metal加速

安装tensorflow-metal插件(仅支持TF 2.5+),利用Apple Metal框架提升推理性能,即使AMD GPU也能获得加速:

pip install tensorflow-metal

内容的提问来源于stack exchange,提问作者Raptor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 00:54:33