TensorFlow训练时显存占用但CPU负载高,是否真的在使用GPU计算?
问题背景
我正在学习神经网络,尝试使用GPU运行训练任务,所用环境如下:
- Python 3.8
- tensorflow-gpu 2.6.0
- PyCharm
- PyCharm的Jupyter插件
- NVIDIA 3080 TI 12GB显卡
已安装CUDA 11.4和CudNN v8.2.4.15,初始导入代码如下:
import os os.environ['TF_FORCE_GPU_ALLOW_GROWTH'] = 'true' import pandas as pd import tensorflow as tf import numpy as np import matplotlib.pyplot as plt from tensorflow import keras from tensorflow.keras.optimizers import Adam from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense, Flatten, Conv1D from tensorflow.keras.preprocessing.sequence import TimeseriesGenerator from sklearn.preprocessing import MinMaxScaler, OneHotEncoder
遇到的问题
启动神经网络训练后任务管理器资源占用异常,看起来在用CPU而非GPU:进程启动后CPU使用率从5%升到40%,显存占用从2GB升到6GB,似乎只占了显卡显存,实际计算由CPU执行,这种情况是否可能?
补充信息1:Jupyter运行日志
"E:\Program Files\JetBrains\PyCharm 2020.3\bin\runnerw.exe" C:\Users\levsh\AppData\Local\Programs\Python\Python38\python.exe -m jupyter notebook --no-browser --notebook-dir=C:/Users/levsh/PycharmProjects/ipynb [W 2021-11-01 02:53:42.606 LabApp] 'notebook_dir' has moved from NotebookApp to ServerApp. This config will be passed to ServerApp. Be sure to update your config before our next release. [W 2021-11-01 02:53:42.607 LabApp] 'notebook_dir' has moved from NotebookApp to ServerApp. This config will be passed to ServerApp. Be sure to update your config before our next release. [W 2021-11-01 02:53:42.607 LabApp] 'notebook_dir' has moved from NotebookApp to ServerApp. This config will be passed to ServerApp. Be sure to update your config before our next release. [I 2021-11-01 02:53:42.614 LabApp] JupyterLab extension loaded from c:\users\levsh\appdata\local\programs\python\python38\lib\site-packages\jupyterlab [I 2021-11-01 02:53:42.614 LabApp] JupyterLab application directory is C:\Users\levsh\AppData\Local\Programs\Python\Python38\share\jupyter\lab [I 02:53:42.620 NotebookApp] Serving notebooks from local directory: C:/Users/levsh/PycharmProjects/ipynb [I 02:53:42.620 NotebookApp] Jupyter Notebook 6.4.3 is running at: [I 02:53:42.620 NotebookApp] http://localhost:8888/?token=a34f13cb67f5cf89bff0a8b8242b69a8727197a98ddb298f [I 02:53:42.620 NotebookApp] or http://127.0.0.1:8888/?token=a34f13cb67f5cf89bff0a8b8242b69a8727197a98ddb298f [I 02:53:42.620 NotebookApp] Use Control-C to stop this server and shut down all kernels (twice to skip confirmation). [C 02:53:42.625 NotebookApp] To access the notebook, open this file in a browser: file:///C:/Users/levsh/AppData/Roaming/jupyter/runtime/nbserver-9860-open.html Or copy and paste one of these URLs: http://localhost:8888/?token=a34f13cb67f5cf89bff0a8b8242b69a8727197a98ddb298f or http://127.0.0.1:8888/?token=a34f13cb67f5cf89bff0a8b8242b69a8727197a98ddb298f [I 02:53:42.626 NotebookApp] 302 GET /api/kernelspecs/ (127.0.0.1) 0.000000ms [I 02:53:42.694 NotebookApp] Kernel started: 036ed1a8-7e7e-49bb-8c86-3efcfb4b993c, name: python3 [W 02:53:42.703 NotebookApp] No session ID specified 2021-11-01 02:53:49.739123: I tensorflow/core/platform/cpu_feature_guard.cc:142] This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN) to use the following CPU instructions in performance-critical operations: AVX AVX2 To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags. 2021-11-01 02:53:50.367361: W tensorflow/core/common_runtime/gpu/gpu_bfc_allocator.cc:39] Overriding allow_growth setting because the TF_FORCE_GPU_ALLOW_GROWTH environment variable is set. Original config value was 0. 2021-11-01 02:53:50.367428: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1510] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 9440 MB memory: -> device: 0, name: NVIDIA GeForce RTX 3080 Ti, pci bus id: 0000:1f:00.0, compute capability: 8.6 2021-11-01 02:53:50.526516: I tensorflow/compiler/mlir/mlir_graph_optimization_pass.cc:185] None of the MLIR Optimization Passes are enabled (registered 2) 2021-11-01 02:53:52.226233: I tensorflow/stream_executor/cuda/cuda_dnn.cc:369] Loaded cuDNN version 8204 2021-11-01 02:53:55.133752: I tensorflow/stream_executor/cuda/cuda_blas.cc:1760] TensorFloat-32 will be used for the matrix multiplication. This will only be logged once.
补充信息2:GPU检测结果
执行GPU设备检测代码如下,日志显示GPU识别正常:
tf.debugging.set_log_device_placement(True) print("Num GPUs Available: ", len(tf.config.list_physical_devices('GPU'))) # Num GPUs Available: 1 tf.config.list_physical_devices('GPU') # [PhysicalDevice(name='/physical_device:GPU:0', device_type='GPU')] sess = tf.compat.v1.Session(config=tf.compat.v1.ConfigProto(log_device_placement=True)) # Device mapping: # /job:localhost/replica:0/task:0/device:GPU:0 -> device: 0, name: NVIDIA GeForce RTX 3080 Ti, pci bus id: 0000:1f:00.0, compute capability: 8.6
对应运行日志:
2021-11-01 04:34:49.606386: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1510] Created device /device:GPU:0 with 9440 MB memory: -> device: 0, name: NVIDIA GeForce RTX 3080 Ti, pci bus id: 0000:1f:00.0, compute capability: 8.6 2021-11-01 04:36:58.513741: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1510] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 9440 MB memory: -> device: 0, name: NVIDIA GeForce RTX 3080 Ti, pci bus id: 0000:1f:00.0, compute capability: 8.6
解答
这种情况基本不可能出现,你观察到的是训练时的正常表现:
- 从你的日志可以明确看到,TensorFlow已经成功识别到3080Ti显卡,加载了cuDNN、启用了TensorFloat-32加速,所有计算操作默认都会调度到GPU执行,显存被占用就说明模型和数据已经加载到GPU侧,不可能出现只占显存不用GPU计算的情况。
- 你看到的CPU占用高是正常现象,CPU此时负责的是数据加载、预处理、批次调度、进度更新等周边工作,尤其是你用到了
TimeseriesGenerator做时序数据生成,这部分逻辑本身就是在CPU上运行的,会占用一定CPU资源。 - Windows任务管理器默认显示的GPU利用率是
3D引擎的占用率,深度学习计算用到的是CUDA/Compute引擎,你需要在任务管理器的GPU性能页,把上方的图表切换到Cuda就能看到真实的GPU计算使用率了。
你可以通过运行大batch的矩阵乘法测试验证GPU利用率,参考代码:
import time import tensorflow as tf # 生成大矩阵 a = tf.random.normal([10000, 10000]) b = tf.random.normal([10000, 10000]) # 预热 _ = tf.matmul(a, b) # 计时测试 start = time.time() for _ in range(10): _ = tf.matmul(a, b) print(f"耗时:{time.time()-start:.2f}s")
运行这段代码时切换到任务管理器的CUDA占用页,就能看到GPU使用率拉满,和CPU跑同规模运算的耗时差非常明显,可以直接确认GPU在正常工作。
内容的提问来源于stack exchange,提问作者Levsha
相关产品推荐
相关产品推荐

