You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何跟踪深度学习中CPU与GPU的运行时间占比?

跟踪深度学习任务中CPU与GPU耗时的方法

我来分享几个实用的方案,不管是通用场景还是你给出的Keras多GPU示例,都能清晰追踪CPU和GPU的耗时占比~

一、通用追踪方案

1. 系统级监控工具

这类工具适合快速查看实时资源使用情况,若要统计耗时占比,可以配合脚本定时采集数据:

  • GPU监控:使用nvidia-smi命令,加上--query-gpu=utilization.gpu,utilization.memory,timestamp --format=csv可以定时输出GPU使用率和时间戳,方便后续统计;也可以用watch -n 1 nvidia-smi实时刷新查看。
  • CPU监控:用top或htop查看CPU整体负载,若要精准统计进程级的CPU耗时,可使用ps -p <进程ID> -o %cpu,etime,或者用Python的psutil库编写脚本,定时获取目标进程的CPU使用率和运行时间。
  • 注意:系统级工具是宏观统计,无法区分代码中具体操作的CPU/GPU耗时,适合整体占比分析。

2. 框架内置性能分析工具

深度学习框架大多自带专业的性能分析工具,能精准追踪每一步操作的设备耗时:

  • TensorFlow/Keras:使用tf.profiler,可以配置追踪CPU和GPU的每一步运算耗时,生成详细的性能报告。
  • PyTorch:使用torch.profiler,支持记录CPU和GPU的操作时间、内存使用等,还能可视化分析。
    这类工具的优势是能定位到具体代码块的耗时,适合优化特定环节。

3. 手动计时(灵活可控)

如果只需要统计关键代码段的耗时,可以手动添加计时逻辑,但要注意GPU异步执行的问题——GPU操作是异步的,直接计时可能得到的是CPU发起请求的时间,而非GPU实际执行时间,需要手动同步:

  • CPU计时:直接用Python的time.time()或time.perf_counter()包裹CPU代码块。
  • GPU计时:在TensorFlow中可以用tf.test.Benchmark()或者手动添加tf.keras.backend.get_session().run(tf.local_variables_initializer())同步GPU操作后再计时。

二、针对你的Keras多GPU示例的具体实现

下面是修改后的代码,结合tf.profiler和手动计时,实现CPU与GPU的耗时追踪:

import tensorflow as tf
from tensorflow.keras.applications import Xception
from tensorflow.keras.utils import multi_gpu_model
import numpy as np
import time

# 配置tf.profiler,追踪CPU和GPU耗时
tf.profiler.experimental.server.start(6009)  # 启动profiler服务,可通过TensorBoard查看
options = tf.profiler.experimental.ProfilerOptions(
    host_tracer_level=2,
    python_tracer_level=2,
    device_tracer_level=1
)

num_samples = 1000
height = 224
width = 224
num_classes = 1000

# 生成模拟数据
x = np.random.random((num_samples, height, width, 3))
y = np.random.random((num_samples, num_classes))

# 实例化基础模型(CPU操作为主)
with tf.profiler.experimental.Trace('CPU_model_init', options=options):
    start_cpu = time.perf_counter()
    base_model = Xception(weights=None,
                          input_shape=(height, width, 3),
                          classes=num_classes)
    end_cpu = time.perf_counter()
    print(f"CPU初始化模型耗时: {end_cpu - start_cpu:.2f}秒")

# 多GPU模型构建与编译(GPU操作为主)
with tf.profiler.experimental.Trace('GPU_model_build', options=options):
    start_gpu = time.perf_counter()
    # 同步GPU,确保计时准确
    tf.keras.backend.get_session().run(tf.local_variables_initializer())
    parallel_model = multi_gpu_model(base_model, gpus=2)
    parallel_model.compile(loss='categorical_crossentropy',
                          optimizer='rmsprop')
    # 再次同步GPU,等待模型构建完成
    tf.keras.backend.get_session().run(tf.local_variables_initializer())
    end_gpu = time.perf_counter()
    print(f"GPU构建多GPU模型耗时: {end_gpu - start_gpu:.2f}秒")

# 训练阶段计时(CPU+GPU混合操作)
with tf.profiler.experimental.Trace('Training', options=options):
    start_train = time.perf_counter()
    # 训练前同步GPU
    tf.keras.backend.get_session().run(tf.local_variables_initializer())
    parallel_model.fit(x, y, epochs=2, batch_size=32)
    # 训练后同步GPU,确保所有操作完成
    tf.keras.backend.get_session().run(tf.local_variables_initializer())
    end_train = time.perf_counter()
    print(f"整体训练耗时: {end_train - start_train:.2f}秒")

# 停止profiler
tf.profiler.experimental.stop()

补充说明:

  1. tf.profiler可视化:运行代码后,启动TensorBoard(tensorboard --logdir=./profile),可以在界面中查看CPU和GPU的详细耗时分布,包括每个操作的设备、时间占比等。
  2. 手动计时注意点:因为GPU操作是异步的,所以在计时前后必须添加GPU同步操作,否则得到的时间会远小于实际GPU执行时间。
  3. 多GPU场景:multi_gpu_model会自动将模型拆分到多个GPU上,profiler会分别追踪每个GPU的耗时,方便分析负载均衡情况。

内容的提问来源于stack exchange,提问作者dward4

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:24:42