You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何准确测量TensorFlow模型加载与推理时的内存消耗?

解决TensorFlow模型加载与推理内存测量问题

为什么你的测量结果每次都不一样?

用psutil的rss测量进程内存出现波动,核心原因有三点:

  • Python垃圾回收时机不确定:测量初始内存m1时,进程中可能残留未回收的临时对象,导致初始值不稳定。
  • TensorFlow内存预分配机制:TensorFlow默认会预分配部分CPU/GPU内存,分配时机可能与模型加载重叠,造成测量差值波动。
  • 模型惰性初始化:部分模型层或参数会在首次使用时才完成初始化,加载阶段并未完全占用所有内存。

更准确的模型加载内存测量方法

方法1:优化psutil测量流程

先稳定进程内存状态,再进行测量:

import os
import psutil
import tensorflow as tf
import gc

# 触发垃圾回收,清理临时内存
gc.collect()
# 初始化TensorFlow,避免加载时的初始化内存干扰
tf.random.normal((1, 1))

process = psutil.Process(os.getpid())
m1 = process.memory_info().rss  # 初始内存
imported = tf.saved_model.load('./predict/001')
print(list(imported.signatures.keys()))
# 再次清理垃圾回收,确保仅统计模型加载的内存
gc.collect()
m2 = process.memory_info().rss  # 加载后内存
print(convert_size(m2 - m1))

方法2:使用TensorFlow内置内存分析工具

用tf.profiler精准追踪TensorFlow内部的内存消耗:

import tensorflow as tf

# 开启内存追踪服务(可通过localhost:6009查看可视化结果)
tf.profiler.experimental.server.start(6009)
# 加载模型
imported = tf.saved_model.load('./predict/001')
# 统计模型参数内存占用
stats = tf.profiler.experimental.profile(
    logdir='./profile',
    cmd='scope',
    options=tf.profiler.experimental.ProfileOptionBuilder.trainable_variables_parameter()
)
# 假设参数为float32(每个占4字节),转换为MB
print(f"模型参数占用内存:{stats.total_parameters * 4 / 1024 / 1024:.2f} MB")

如何测量推理时的最大内存?

模型加载内存仅为静态参数占用,推理时会生成大量中间张量(如层输出、激活值),峰值内存通常远高于加载内存。可通过以下方法测量:

方法1:用psutil追踪推理全程峰值内存

import os
import psutil
import tensorflow as tf
import gc
import numpy as np

gc.collect()
tf.random.normal((1, 1))
process = psutil.Process(os.getpid())

# 加载模型
imported = tf.saved_model.load('./predict/001')
infer = imported.signatures['serving_default']  # 替换为你的模型签名key

# 准备与模型输入形状匹配的测试数据
test_input = tf.random.normal((1, 224, 224, 3))  # 根据实际模型调整

max_memory = 0
# 多次推理取稳定峰值,避免首次推理的初始化干扰
for _ in range(5):
    gc.collect()
    pre_infer_mem = process.memory_info().rss
    # 执行推理
    result = infer(test_input)
    post_infer_mem = process.memory_info().rss
    # 更新峰值内存
    current_peak = max(pre_infer_mem, post_infer_mem)
    if current_peak > max_memory:
        max_memory = current_peak

# 计算推理相对于模型加载后的额外内存消耗
post_load_mem = process.memory_info().rss
print(f"推理峰值内存:{convert_size(max_memory - post_load_mem)}")

方法2:GPU内存精准监控(适用于GPU推理)

开启内存增长模式,用TensorFlow内置工具查看实时GPU内存:

import tensorflow as tf

# 启用GPU内存增长,避免预分配过多内存
gpus = tf.config.list_physical_devices('GPU')
if gpus:
    try:
        for gpu in gpus:
            tf.config.experimental.set_memory_growth(gpu, True)
    except RuntimeError as e:
        print(e)

# 加载模型
imported = tf.saved_model.load('./predict/001')
infer = imported.signatures['serving_default']

# 查看模型加载后的GPU已用内存
load_memory = tf.config.experimental.get_memory_info('GPU:0')['used']

# 执行推理
test_input = tf.random.normal((1, 224, 224, 3))
result = infer(test_input)

# 查看推理后的GPU已用内存
infer_memory = tf.config.experimental.get_memory_info('GPU:0')['used']
print(f"推理额外占用GPU内存:{(infer_memory - load_memory)/1024/1024:.2f} MB")

内容的提问来源于stack exchange,提问作者SaD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 15:00:59