Windows环境下TensorFlow 1.8 GPU版本未启用GPU加速问题求助
我在使用TensorFlow的AlexNet卷积神经网络训练模型时,遇到了一个奇怪的问题——训练过程没有报错,TensorBoard能正常访问,模型测试也没问题,但GPU使用率基本维持在0%,偶尔才短暂跳到25%,反而CPU占用率高达90%以上,看起来模型根本没在GPU上跑,而是用了CPU。
我的环境配置
Windows 8.1 x64 GPU 1070 driver version 3.88 tensorflow-gpu 1.8.0 CUDA toolkit v9.0 cuDNN version 7
已完成的正确性验证
我已经正常导入TensorFlow,并且做了三个测试来确认安装没问题:
测试1:查看本地设备
运行以下代码:
from tensorflow.python.client import device_lib print(device_lib.list_local_devices())
输出结果:
[name: "/device:CPU:0" device_type: "CPU" memory_limit: 268435456 locality { } incarnation: 625346735515728619 , name: "/device:GPU:0" device_type: "GPU" memory_limit: 6911164212 locality { bus_id: 1 links { } } incarnation: 15764160474642097170 physical_device_desc: "device: 0, name: GeForce GTX 1070, pci bus id: 0000:01:00.0, compute capability: 6.1" ]
测试2:基础TensorFlow会话测试
代码:
import tensorflow as tf hello = tf.constant('Hello, TensorFlow!') sess = tf.Session() print(sess.run(hello))
输出:
b'Hello, TensorFlow!'
测试3:TensorFlow GPU支持检测脚本
代码:
import ctypes import imp import sys def main(): try: import tensorflow as tf print("TensorFlow successfully installed.") if tf.test.is_built_with_cuda(): print("The installed version of TensorFlow includes GPU support.") else: print("The installed version of TensorFlow does not include GPU support.") sys.exit(0) except ImportError: print("ERROR: Failed to import the TensorFlow module.") candidate_explanation = False python_version = sys.version_info.major, sys.version_info.minor print("\n- Python version is %d.%d." % python_version) if not (python_version == (3, 5) or python_version == (3, 6)): candidate_explanation = True print("- The official distribution of TensorFlow for Windows requires " "Python version 3.5 or 3.6.") try: _, pathname, _ = imp.find_module("tensorflow") print("\n- TensorFlow is installed at: %s" % pathname) except ImportError: candidate_explanation = False print(""" - No module named TensorFlow is installed in this Python environment. You may install it using the command `pip install tensorflow`.""") try: msvcp140 = ctypes.WinDLL("msvcp140.dll") except OSError: candidate_explanation = True print(""" - Could not load 'msvcp140.dll'. TensorFlow requires that this DLL be installed in a directory that is named in your %PATH% environment variable. You may install this DLL by downloading Microsoft Visual C++ 2015 Redistributable Update 3 from this URL: https://www.microsoft.com/en-us/download/details.aspx?id=53587""") try: cudart64_80 = ctypes.WinDLL("cudart64_80.dll") except OSError: candidate_explanation = True print(""" - Could not load 'cudart64_80.dll'. The GPU version of TensorFlow requires that this DLL be installed in a directory that is named in your %PATH% environment variable. Download and install CUDA 8.0 from this URL: https://developer.nvidia.com/cuda-toolkit""") try: nvcuda = ctypes.WinDLL("nvcuda.dll") except OSError: candidate_explanation = True print(""" - Could not load 'nvcuda.dll'. The GPU version of TensorFlow requires that this DLL be installed in a directory that is named in your %PATH% environment variable. Typically it is installed in 'C:\\Windows\\System32'. If it is not present, ensure that you have a CUDA-capable GPU with the correct driver installed.""") cudnn5_found = False try: cudnn5 = ctypes.WinDLL("cudnn64_5.dll") cudnn5_found = True except OSError: candidate_explanation = True print(""" - Could not load 'cudnn64_5.dll'. The GPU version of TensorFlow requires that this DLL be installed in a directory that is named in your %PATH% environment variable. Note that installing cuDNN is a separate step from installing CUDA, and it is often found in a different directory from the CUDA DLLs. You may install the necessary DLL by downloading cuDNN 5.1 from this URL: https://developer.nvidia.com/cudnn""") cudnn6_found = False try: cudnn = ctypes.WinDLL("cudnn64_6.dll") cudnn6_found = True except OSError: candidate_explanation = True if not cudnn5_found or not cudnn6_found: print() if not cudnn5_found and not cudnn6_found: print("- Could not find cuDNN.") elif not cudnn5_found: print("- Could not find cuDNN 5.1.") else: print("- Could not find cuDNN 6.") print(""" The GPU version of TensorFlow requires that the correct cuDNN DLL be installed in a directory that is named in your %PATH% environment variable. Note that installing cuDNN is a separate step from installing CUDA, and it is often found in a different directory from the CUDA DLLs. The correct version of cuDNN depends on your version of TensorFlow: * TensorFlow 1.2.1 or earlier requires cuDNN 5.1. ('cudnn64_5.dll') * TensorFlow 1.3 or later requires cuDNN 6. ('cu""") ## 可能的排查方向与解决方案 既然环境验证都通过了,那问题大概率出在**模型代码的设备指定**或者**TensorFlow的默认设备分配逻辑**上,给你几个实操性的排查点: - **强制绑定GPU设备运行** TensorFlow默认会优先用GPU,但有时候可能因为某些操作意外落到CPU上(虽然AlexNet的核心操作都支持GPU),或者代码里隐性指定了CPU。你可以把模型构建、训练的所有代码放到GPU设备上下文里: ```python import tensorflow as tf # 强制使用第一块GPU with tf.device('/device:GPU:0'): # 这里放你的AlexNet模型定义、数据集加载、训练循环代码 pass
如果模型里有不支持GPU的操作,这段代码会直接报错,帮你快速定位问题环节。
排查数据加载是否成为瓶颈
CPU满负荷、GPU闲得慌,最常见的原因就是数据预处理/加载速度跟不上GPU的计算速度——GPU一直在等CPU喂数据,自然使用率上不去。可以试试:- 改用
tf.data.DatasetAPI代替传统的numpy数组喂数,它支持多线程异步预处理和预加载,能有效缓解CPU瓶颈 - 把数据预处理逻辑(比如图片缩放、归一化)尽量放到TensorFlow图中(用
tf.image系列函数),而不是在Python层面用numpy处理,这样预处理可以在GPU上并行执行
- 改用
显式配置会话的GPU参数
虽然默认会话会自动使用GPU,但显式配置可以确保GPU被正确调用,还能排查设备分配问题:config = tf.ConfigProto() config.gpu_options.allow_growth = True # 按需分配GPU内存,避免一次性占满 config.log_device_placement = True # 打印每个操作的设备分配情况 sess = tf.Session(config=config)开启
log_device_placement后,训练时会输出每个层/操作被分配到的设备,一目了然是不是真的在GPU上运行。调大训练的batch size
如果你的batch size设置得太小(比如个位数),GPU可能刚启动计算就完成了一轮,导致使用率上不去。可以尝试调大batch size(比如64、128,根据你的GPU内存调整),看看GPU使用率会不会提升。再次确认CUDA与cuDNN的路径配置
你的版本匹配是对的(TensorFlow 1.8.0对应CUDA 9.0+cuDNN7),但要确保cuDNN的三个核心文件(cudnn64_7.dll、cudnn.h、cudnn.lib)已经放到CUDA的对应目录里(比如C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v9.0\bin、include、lib\x64),并且这些路径都在系统的PATH环境变量中。
内容的提问来源于stack exchange,提问作者jayinbluecity

