TensorFlow与任意CUDA版本搭配调用GPU接口时触发RuntimeError: No allocator statistics错误求助
问题描述
自昨日起,使用TensorFlow时遇到异常:无论搭配哪个版本的CUDA,执行import tensorflow as tf; tf.test.is_gpu_available()就会触发错误,调用Convolutional()函数时也会出现相同错误。执行日志如下:
>>> import tensorflow as tf 2021-04-16 21:23:14.876381: I tensorflow/stream_executor/platform/default/dso_loader.cc:49] Successfully opened dynamic library cudart64_110.dll INFO:tensorflow:Enabling eager execution INFO:tensorflow:Enabling v2 tensorshape INFO:tensorflow:Enabling resource variables INFO:tensorflow:Enabling tensor equality INFO:tensorflow:Enabling control flow v2 --> tf.test.is_gpu_available() WARNING:tensorflow:From <stdin>:1: is_gpu_available (from tensorflow.python.framework.test_util) is deprecated and will be removed in a future version. Instructions for updating: Use tf.config.list_physical_devices('GPU') instead. 2021-04-16 21:23:25.481405: I tensorflow/core/platform/cpu_feature_guard.cc:142] This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN) to use the following CPU instructions in performance-critical operations: AVX AVX2 To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags. 2021-04-16 21:23:25.484960: I tensorflow/stream_executor/platform/default/dso_loader.cc:49] Successfully opened dynamic library nvcuda.dll 2021-04-16 21:23:25.520614: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1733] Found device 0 with properties: pciBusID: 0000:09:00.0 name: NVIDIA GeForce RTX 2060 SUPER computeCapability: 7.5 coreClock: 1.65GHz coreCount: 34 deviceMemorySize: 8.00GiB deviceMemoryBandwidth: 417.29GiB/s 2021-04-16 21:23:25.520759: I tensorflow/stream_executor/platform/default/dso_loader.cc:49] Successfully opened dynamic library cudart64_110.dll 2021-04-16 21:23:25.529842: I tensorflow/stream_executor/platform/default/dso_loader.cc:49] Successfully opened dynamic library cublas64_11.dll 2021-04-16 21:23:25.529971: I tensorflow/stream_executor/platform/default/dso_loader.cc:49] Successfully opened dynamic library cublasLt64_11.dll 2021-04-16 21:23:25.534167: I tensorflow/stream_executor/platform/default/dso_loader.cc:49] Successfully opened dynamic library cufft64_10.dll 2021-04-16 21:23:25.535573: I tensorflow/stream_executor/platform/default/dso_loader.cc:49] Successfully opened dynamic library curand64_10.dll 2021-04-16 21:23:25.540947: I tensorflow/stream_executor/platform/default/dso_loader.cc:49] Successfully opened dynamic library cusolver64_11.dll 2021-04-16 21:23:25.544673: I tensorflow/stream_executor/platform/default/dso_loader.cc:49] Successfully opened dynamic library cusparse64_11.dll 2021-04-16 21:23:25.545533: I tensorflow/stream_executor/platform/default/dso_loader.cc:49] Successfully opened dynamic library cudnn64_8.dll 2021-04-16 21:23:25.545662: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1871] Adding visible gpu devices: 0 2021-04-16 21:23:26.045993: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1258] Device interconnect StreamExecutor with strength 1 edge matrix: 2021-04-16 21:23:26.046109: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1264] 0 2021-04-16 21:23:26.047907: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1277] 0: N 2021-04-16 21:23:26.048739: I tensorflow/core/common_runtime/gpu/gpu_process_state.cc:210] Using CUDA malloc Async allocator for GPU. Traceback (most recent call last): File "<stdin>", line 1, in <module> File "C:\Users\Weise\AppData\Roaming\Python\Python39\site-packages\tensorflow\python\util\deprecation.py", line 337, in new_func return func(*args, **kwargs) File "C:\Users\Weise\AppData\Roaming\Python\Python39\site-packages\tensorflow\python\framework\test_util.py", line 1600, in is_gpu_available for local_device in device_lib.list_local_devices(): File "C:\Users\Weise\AppData\Roaming\Python\Python39\site-packages\tensorflow\python\client\device_lib.py", line 43, in list_local_devices _convert(s) for s in _pywrap_device_lib.list_devices(serialized_config) RuntimeError: No allocator statistics
排查建议
我之前遇到过类似的报错,结合你的日志信息,给你几个可行的排查方向:
替换弃用的检测函数:日志里已经明确提示
is_gpu_available()已被弃用,建议改用官方推荐的tf.config.list_physical_devices('GPU')来检测GPU,代码如下:import tensorflow as tf print(tf.config.list_physical_devices('GPU'))这个函数更稳定,也能准确识别GPU设备状态。
禁用CUDA异步内存分配器:日志显示TensorFlow正在使用
CUDA malloc Async allocator,这个分配器可能和你的环境存在兼容性问题。你可以通过设置环境变量切换到默认分配器:
Windows系统下,启动Python前执行:set TF_GPU_ALLOCATOR=default或者在代码开头添加:
import os os.environ['TF_GPU_ALLOCATOR'] = 'default' import tensorflow as tf严格匹配版本兼容性:虽然你尝试了多个CUDA版本,但要确保TensorFlow、CUDA、cuDNN和Python的版本完全匹配。比如你日志里加载的是
cudart64_110.dll(对应CUDA11.0)和cudnn64_8.dll,那么TensorFlow应该使用2.4.x版本,同时Python版本要在3.6-3.9之间(TF2.4支持Python3.9),版本不匹配很容易引发这类底层错误。清理并重新安装TensorFlow:有时候缓存文件损坏会导致奇怪的问题,建议先卸载当前的TensorFlow:
pip uninstall tensorflow -y然后清理Python缓存目录(比如
C:\Users\Weise\AppData\Roaming\Python\Python39\site-packages下的tensorflow相关文件夹),再重新安装对应版本的TensorFlow:pip install tensorflow==2.4.0开启GPU内存增长模式:尝试启用GPU内存动态增长,避免一次性占用过多内存导致的分配问题,代码如下:
import tensorflow as tf gpus = tf.config.list_physical_devices('GPU') if gpus: try: for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True) logical_gpus = tf.config.list_logical_devices('GPU') print(len(gpus), "Physical GPUs,", len(logical_gpus), "Logical GPUs") except RuntimeError as e: print(e)
内容的提问来源于stack exchange,提问作者weiserhase

