GPU可在PyTorch中运行但TensorFlow无法调用的问题求助
TensorFlow无法调用GPU问题排查
环境配置
当前conda环境的关键包版本:
(myen2v) C:\Users\Jan>conda list cudnn # packages in environment at D:\BitDownlD\Anaconda8\envs\myen2v: # # Name Version Build Channel cudnn 8.9.2.26 cuda11_0 anaconda (myen2v) C:\Users\Jan>conda list cuda # packages in environment at D:\BitDownlD\Anaconda8\envs\myen2v: # # Name Version Build Channel cudatoolkit 11.8.0 hd77b12b_0 (myen2v) C:\Users\Jan>conda list torch # packages in environment at D:\BitDownlD\Anaconda8\envs\myen2v: # # Name Version Build Channel pytorch 2.0.1 cpu_py38hb0bdfb8_0 torch 2.1.0 pypi_0 pypi (myen2v) C:\Users\Jan>conda list tensor # packages in environment at D:\BitDownlD\Anaconda8\envs\myen2v: # # Name Version Build Channel tensorboard 2.13.0 pypi_0 pypi tensorboard-data-server 0.7.1 pypi_0 pypi tensorboard-plugin-wit 1.8.1 py38haa95532_0 tensorflow 2.13.0 pypi_0 pypi tensorflow-base 2.3.0 eigen_py38h75a453f_0 tensorflow-estimator 2.13.0 pypi_0 pypi tensorflow-gpu 2.3.0 pypi_0 pypi tensorflow-gpu-estimator 2.3.0 pypi_0 pypi tensorflow-io-gcs-filesystem 0.31.0 pypi_0 pypi
PyTorch GPU运行正常
执行以下代码可成功在GPU上运行:
# Create tensors on GPU a = torch.tensor([1, 2, 3], device="cuda") b = torch.tensor([4, 5, 6], device="cuda") # Perform operations on GPU c = a + b print(c)
输出:
tensor([5, 7, 9], device='cuda:0')
TensorFlow运行报错
测试代码1及报错
import tensorflow as tf physical_devices = tf.config.list_physical_devices('GPU') print("Num GPUs:", len(physical_devices))
报错信息:
--------------------------------------------------------------------------- AttributeError Traceback (most recent call last) Cell In[12], line 1 ----> 1 physical_devices = tf.config.list_physical_devices('GPU') 2 print("Num GPUs:", len(physical_devices)) AttributeError: module 'tensorflow' has no attribute 'config'
测试代码2及报错
import tensorflow as tf print("Num of GPUs available: ", len(tf.test.gpu_device_name()))
报错信息:
--------------------------------------------------------------------------- AttributeError Traceback (most recent call > last) Cell In[13], line 2 > 1 import tensorflow as tf > ----> 2 print("Num of GPUs available: ", len(tf.test.gpu_device_name())) > > AttributeError: module 'tensorflow' has no attribute 'test'
问题原因及解决步骤
问题原因
环境中同时安装了多个版本的TensorFlow包(2.13.0和2.3.0),且混装了tensorflow、tensorflow-base、tensorflow-gpu,导致包冲突,TensorFlow模块加载异常。
解决步骤
卸载所有TensorFlow相关包
在conda环境中执行以下命令:pip uninstall tensorflow tensorflow-base tensorflow-gpu tensorflow-estimator tensorflow-gpu-estimator tensorflow-io-gcs-filesystem若部分包由conda安装,可改用conda卸载:
conda remove tensorflow tensorflow-base tensorflow-gpu tensorflow-estimator tensorflow-gpu-estimator tensorflow-io-gcs-filesystem安装与CUDA版本兼容的TensorFlow
当前CUDA版本为11.8,TensorFlow 2.13.0官方支持CUDA 11.8,直接安装主包即可(TensorFlow 2.x之后无需单独安装tensorflow-gpu,主包已包含GPU支持):pip install tensorflow==2.13.0验证安装
执行以下代码验证GPU是否可用:import tensorflow as tf print("Num GPUs Available: ", len(tf.config.list_physical_devices('GPU'))) # 进一步验证GPU计算 tf.debugging.set_log_device_placement(True) a = tf.constant([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]]) b = tf.constant([[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]]) c = tf.matmul(a, b) print(c)若输出中包含GPU设备信息,则说明配置成功。
内容的提问来源于stack exchange,提问作者user22758952
相关产品推荐
相关产品推荐

