TensorFlow 2.8.0在Windows平台GeForce GTX 1650 Ti GPU上model.fit()停滞于Epoch 1的问题求助
TensorFlow 2.8.0在GTX 1650 Ti上训练时无限停滞在Epoch 1并崩溃
我正尝试在Windows系统的GeForce GTX 1650 Ti GPU上运行TensorFlow 2.8.0,但遇到了非常棘手的问题:调用model.fit()时,任何模型都会无限期卡在Epoch 1,直到内核崩溃重启(已经在Jupyter Notebook和Spyder里都测试过了)。
根据TensorFlow官方文档,我已经下载并安装了对应的CUDA和cuDNN版本,并且完成了以下验证(同时确认了TensorFlow能检测到GPU):
已完成的环境验证
CUDA版本确认(要求为11.2)
- 命令行执行:
nvcc --version
输出:Build cuda_11.2.r11.2/compiler.29373293_0 - Python环境执行:
输出:import tensorflow.python.platform.build_info as build print(build.build_info['cuda_version'])'64_112'
cuDNN版本确认(要求为8.1)
Python环境执行:
import tensorflow.python.platform.build_info as build print(build.build_info['cudnn_version'])
输出:'64_8'
注:实际安装的是适配CUDA 11.0/11.1/11.2的cuDNN v8.1.1(2021年2月26日发布),推测此输出显示为v8属于正常情况。
GPU检测结果
- 执行
tf.config.list_physical_devices('GPU'),输出:[PhysicalDevice(name='/physical_device:GPU:0', device_type='GPU')] - 执行
tf.test.is_gpu_available(),输出:True - 执行
tf.test.gpu_device_name(),输出:This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN) to use the following CPU instructions in performance-critical operations: AVX AVX2 To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags. Created device /device:GPU:0 with 2153 MB memory: -> device: 0, name: NVIDIA GeForce GTX 1650 Ti, pci bus id: 0000:01:00.0, compute capability: 7.5
特殊情况说明
令人困惑的是,除了某段Stack Overflow问答中的代码可以正常运行外,包括TensorFlow官方CNN教程在内的其他代码均无法正常执行。我在过去数小时内尝试了各类TensorFlow代码,唯有该StackOverflow代码不会在Epoch 1停滞。
另外,我尝试通过os.environ['CUDA_VISIBLE_DEVICES'] = '-1'强制仅使用CPU运行,此时所有代码均能正常工作。
恳请各位帮忙排查并解决这个问题!
内容的提问来源于stack exchange,提问作者Anselm
相关产品推荐
相关产品推荐

