定义tf.distribute.MirroredStrategy后GPU显存几乎占满问题咨询
TensorFlow初始化MirroredStrategy后GPU显存被占满问题
问题表现
使用TensorFlow 2.9.1开发时,完成分布式训练策略定义后,GPU显存几乎被完全占满,仅需两行代码即可复现:
import tensorflow as tf strat = tf.distribute.MirroredStrategy()
分阶段显存状态
- 仅执行TensorFlow导入语句后,通过
nvidia-smi查询显卡状态,两块Quadro P6000显存占用均为0MiB,无相关运行进程:
Fri Jun 10 03:01:47 2022 +-----------------------------------------------------------------------------+ | NVIDIA-SMI 470.103.01 Driver Version: 470.103.01 CUDA Version: 11.4 | |-------------------------------+----------------------+----------------------+ | GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |===============================+======================+======================| | 0 Quadro P6000 Off | 00000000:04:00.0 Off | Off | | 26% 25C P8 9W / 250W | 0MiB / 24449MiB | 0% Default | | | | N/A | +-------------------------------+----------------------+----------------------+ | 1 Quadro P6000 Off | 00000000:06:00.0 Off | Off | | 26% 20C P8 7W / 250W | 0MiB / 24449MiB | 0% Default | | | | N/A | +-------------------------------+----------------------+----------------------+ +-----------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=============================================================================| | No running processes found | +-----------------------------------------------------------------------------+
- 执行
MirroredStrategy初始化语句后,再次查询显卡状态,两块GPU显存占用均达到23951MiB / 24449MiB,接近单卡总显存容量,对应Python进程单卡显存占用达23949MiB,此时GPU利用率为0%:
Fri Jun 10 03:02:43 2022 +-----------------------------------------------------------------------------+ | NVIDIA-SMI 470.103.01 Driver Version: 470.103.01 CUDA Version: 11.4 | |-------------------------------+----------------------+----------------------+ | GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |===============================+======================+======================| | 0 Quadro P6000 Off | 00000000:04:00.0 Off | Off | | 26% 29C P0 59W / 250W | 23951MiB / 24449MiB | 0% Default | | | | N/A | +-------------------------------+----------------------+----------------------+ | 1 Quadro P6000 Off | 00000000:06:00.0 Off | Off | | 26% 25C P0 58W / 250W | 23951MiB / 24449MiB | 0% Default | | | | N/A | +-------------------------------+----------------------+----------------------+ +-----------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=============================================================================| | 0 N/A N/A 1833720 C python 23949MiB | | 1 N/A N/A 1833720 C python 23949MiB | +-----------------------------------------------------------------------------+
- 代码执行过程终端输出日志如下,提示已完成两块GPU设备创建,正在使用MirroredStrategy开展分布式调度:
2022-06-10 03:02:37.442336: I tensorflow/core/platform/cpu_feature_guard.cc:193] This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN) to use the following CPU instructions in performance-critical operations: AVX2 FMA To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags. 2022-06-10 03:02:39.136390: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1532] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 23678 MB memory: -> device: 0, name: Quadro P6000, pci bus id: 0000:04:00.0, compute capability: 6.1 2022-06-10 03:02:39.139204: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1532] Created device /job:localhost/replica:0/task:0/device:GPU:1 with 23678 MB memory: -> device: 1, name: Quadro P6000, pci bus id: 0000:06:00.0, compute capability: 6.1 INFO:tensorflow:Using MirroredStrategy with devices ('/job:localhost/replica:0/task:0/device:GPU:0', '/job:localhost/replica:0/task:0/device:GPU:1')
运行环境
- 操作系统:Linux
- Python版本:3.10.4 [GCC 7.5.0]
- TensorFlow版本:2.9.1
- 依赖版本:CUDA 11.2.2、cuDNN v8.2.1
产生原因
该现象是TensorFlow默认显存分配机制导致,不属于程序错误:
TensorFlow默认会在GPU设备初始化时预占几乎全部可用显存,减少运行过程中动态申请显存产生的性能损耗。初始化MirroredStrategy时,框架会枚举并初始化所有可见的物理GPU设备,直接触发默认的显存预占逻辑,因此哪怕还未启动训练任务、GPU利用率为0,显存也会被提前占满。
解决方法
方法1:开启显存按需增长
在导入TensorFlow后、初始化分布式策略前,配置GPU显存按需申请,不提前预占全部显存,显存占用会随任务运行动态增长:
import tensorflow as tf gpus = tf.config.list_physical_devices('GPU') for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True) strat = tf.distribute.MirroredStrategy()
注意:显存配置必须在GPU设备初始化前完成设置,否则会抛出运行时错误。
方法2:硬限制单卡显存占用上限
如果需要给TensorFlow进程分配固定大小的显存,避免占用全部显卡资源,可以通过虚拟设备配置设置单卡显存上限,以下示例为每块GPU分配10GB可用显存:
import tensorflow as tf gpus = tf.config.list_physical_devices('GPU') for gpu in gpus: tf.config.set_logical_device_configuration( gpu, [tf.config.LogicalDeviceConfiguration(memory_limit=10240)] ) strat = tf.distribute.MirroredStrategy()
方法3:通过环境变量全局配置
启动脚本前设置环境变量TF_FORCE_GPU_ALLOW_GROWTH=true,等价于全局开启显存按需增长,无需修改业务代码,示例启动命令:
TF_FORCE_GPU_ALLOW_GROWTH=true python your_train_script.py
内容的提问来源于stack exchange,提问作者sandsoft
相关产品推荐
相关产品推荐

