如何跟踪深度学习中CPU与GPU的运行时间占比?
跟踪深度学习任务中CPU与GPU耗时的方法
我来分享几个实用的方案,不管是通用场景还是你给出的Keras多GPU示例,都能清晰追踪CPU和GPU的耗时占比~
一、通用追踪方案
1. 系统级监控工具
这类工具适合快速查看实时资源使用情况,若要统计耗时占比,可以配合脚本定时采集数据:
- GPU监控:使用
nvidia-smi命令,加上--query-gpu=utilization.gpu,utilization.memory,timestamp --format=csv可以定时输出GPU使用率和时间戳,方便后续统计;也可以用watch -n 1 nvidia-smi实时刷新查看。 - CPU监控:用
top或htop查看CPU整体负载,若要精准统计进程级的CPU耗时,可使用ps -p <进程ID> -o %cpu,etime,或者用Python的psutil库编写脚本,定时获取目标进程的CPU使用率和运行时间。 - 注意:系统级工具是宏观统计,无法区分代码中具体操作的CPU/GPU耗时,适合整体占比分析。
2. 框架内置性能分析工具
深度学习框架大多自带专业的性能分析工具,能精准追踪每一步操作的设备耗时:
- TensorFlow/Keras:使用
tf.profiler,可以配置追踪CPU和GPU的每一步运算耗时,生成详细的性能报告。 - PyTorch:使用
torch.profiler,支持记录CPU和GPU的操作时间、内存使用等,还能可视化分析。
这类工具的优势是能定位到具体代码块的耗时,适合优化特定环节。
3. 手动计时(灵活可控)
如果只需要统计关键代码段的耗时,可以手动添加计时逻辑,但要注意GPU异步执行的问题——GPU操作是异步的,直接计时可能得到的是CPU发起请求的时间,而非GPU实际执行时间,需要手动同步:
- CPU计时:直接用Python的
time.time()或time.perf_counter()包裹CPU代码块。 - GPU计时:在TensorFlow中可以用
tf.test.Benchmark()或者手动添加tf.keras.backend.get_session().run(tf.local_variables_initializer())同步GPU操作后再计时。
二、针对你的Keras多GPU示例的具体实现
下面是修改后的代码,结合tf.profiler和手动计时,实现CPU与GPU的耗时追踪:
import tensorflow as tf from tensorflow.keras.applications import Xception from tensorflow.keras.utils import multi_gpu_model import numpy as np import time # 配置tf.profiler,追踪CPU和GPU耗时 tf.profiler.experimental.server.start(6009) # 启动profiler服务,可通过TensorBoard查看 options = tf.profiler.experimental.ProfilerOptions( host_tracer_level=2, python_tracer_level=2, device_tracer_level=1 ) num_samples = 1000 height = 224 width = 224 num_classes = 1000 # 生成模拟数据 x = np.random.random((num_samples, height, width, 3)) y = np.random.random((num_samples, num_classes)) # 实例化基础模型(CPU操作为主) with tf.profiler.experimental.Trace('CPU_model_init', options=options): start_cpu = time.perf_counter() base_model = Xception(weights=None, input_shape=(height, width, 3), classes=num_classes) end_cpu = time.perf_counter() print(f"CPU初始化模型耗时: {end_cpu - start_cpu:.2f}秒") # 多GPU模型构建与编译(GPU操作为主) with tf.profiler.experimental.Trace('GPU_model_build', options=options): start_gpu = time.perf_counter() # 同步GPU,确保计时准确 tf.keras.backend.get_session().run(tf.local_variables_initializer()) parallel_model = multi_gpu_model(base_model, gpus=2) parallel_model.compile(loss='categorical_crossentropy', optimizer='rmsprop') # 再次同步GPU,等待模型构建完成 tf.keras.backend.get_session().run(tf.local_variables_initializer()) end_gpu = time.perf_counter() print(f"GPU构建多GPU模型耗时: {end_gpu - start_gpu:.2f}秒") # 训练阶段计时(CPU+GPU混合操作) with tf.profiler.experimental.Trace('Training', options=options): start_train = time.perf_counter() # 训练前同步GPU tf.keras.backend.get_session().run(tf.local_variables_initializer()) parallel_model.fit(x, y, epochs=2, batch_size=32) # 训练后同步GPU,确保所有操作完成 tf.keras.backend.get_session().run(tf.local_variables_initializer()) end_train = time.perf_counter() print(f"整体训练耗时: {end_train - start_train:.2f}秒") # 停止profiler tf.profiler.experimental.stop()
补充说明:
- tf.profiler可视化:运行代码后,启动TensorBoard(
tensorboard --logdir=./profile),可以在界面中查看CPU和GPU的详细耗时分布,包括每个操作的设备、时间占比等。 - 手动计时注意点:因为GPU操作是异步的,所以在计时前后必须添加GPU同步操作,否则得到的时间会远小于实际GPU执行时间。
- 多GPU场景:
multi_gpu_model会自动将模型拆分到多个GPU上,profiler会分别追踪每个GPU的耗时,方便分析负载均衡情况。
内容的提问来源于stack exchange,提问作者dward4
相关产品推荐
相关产品推荐

