TensorFlow目标检测模型for循环运行内存线性增长问题求解
问题根因
内存线性增长的核心原因是每次循环调用detect函数时,内部的tf.*操作会持续向默认计算图新增节点,手动删除变量、清空session无法销毁这些累积节点,最终导致内存持续上涨。你使用的K.clear_session属于TF1.x遗留API,在TF2.x eager模式下使用反而会残留无效状态,加剧内存泄漏。
解决方案
核心思路是提前固定计算图逻辑,避免每次推理生成新节点,同时切断不必要的TF张量内存引用:
- 用
tf.function封装推理逻辑,提前固定输入签名,避免重复构建计算图 - 所有后处理操作优先用numpy实现,结果转成普通数值/数组返回,不持有TF张量引用
- 移除推理过程中的
K.clear_session调用,避免频繁重置会话产生无效状态
修改后代码示例
import numpy as np import tensorflow as tf import tensorflow_hub as tfhub # 1. 全局仅执行一次:加载模型、开启GPU内存动态分配 gpus = tf.config.list_physical_devices('GPU') if gpus: try: for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True) except RuntimeError as e: print(e) # 加载模型仅执行一次,不要放到循环或detect函数内 detector = tfhub.load("你的模型本地路径或模型名") # 2. 用tf.function封装推理逻辑,固定输入签名 @tf.function(input_signature=[tf.TensorSpec(shape=[None, None, 3], dtype=tf.uint8)]) def run_inference(img): detector_output = detector(tf.expand_dims(img, axis=0)) classes = detector_output['detection_classes'][0] most_likely = classes[0] box = detector_output['detection_boxes'][0][0] return box, most_likely # 3. detect仅做轻量后处理,全部用numpy运算 def detect(img): height, width = img.shape[:2] box, most_likely = run_inference(img) # 后处理转numpy操作,不生成新的TF计算节点 box = box.numpy() * [height, width, height, width] box = box.astype(np.int16) # 结果转普通数值返回,不持有TF张量引用 return box, most_likely.numpy() # 4. 循环逻辑保持不变 boxes = [] for i in range(num_images): box, m = detect(imgs[i]) boxes.append(box)
注:你原代码中定义的
second_变量未使用,若需要返回该参数可自行添加到返回值列表。
额外优化建议
- 如果你的输入图片尺寸固定,可以把
run_inference的输入签名调整为固定尺寸(比如shape=[640, 640, 3]),进一步提升推理速度 - 所有
tf.*开头的操作、变量定义全部放到循环外部,不要在循环内部新增任何TF原生操作 - 若仍存在小幅内存增长,可每处理100~200张图片手动触发一次python垃圾回收:
import gc; gc.collect()
内容的提问来源于stack exchange,提问作者Olli
相关产品推荐
相关产品推荐

