YOLOv3(Keras) GPU训练遇cuBLAS_STATUS_EXECUTION_FAILED错误求助
问题
环境配置
- 操作系统:Windows 11 Pro
- GPU:NVIDIA GeForce RTX 3070
- Python:3.7.16
- CUDA:10.0.130
- cuDNN:7.6.5
- TensorFlow:1.15.0、tensorflow-gpu 1.15.0
- Keras:2.2.4
报错情况
运行YOLOv3训练脚本时无法使用GPU训练,报错failed to run cuBLAS routine: CUBLAS_STATUS_EXECUTION_FAILED,完整报错输出如下:
Create YOLOv3 model with 9 anchors and 47 classes. C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\saving.py:1140: UserWarning: Skipping loading of weights for layer conv2d_59 due to mismatch in shape ((1, 1, 1024, 156) vs (255, 1024, 1, 1)). weight_values[i].shape)) C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\saving.py:1140: UserWarning: Skipping loading of weights for layer conv2d_59 due to mismatch in shape ((156,) vs (255,)). weight_values[i].shape)) C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\saving.py:1140: UserWarning: Skipping loading of weights for layer conv2d_67 due to mismatch in shape ((1, 1, 512, 156) vs (255, 512, 1, 1)). weight_values[i].shape)) C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\saving.py:1140: UserWarning: Skipping loading of weights for layer conv2d_67 due to mismatch in shape ((156,) vs (255,)). weight_values[i].shape)) C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\saving.py:1140: UserWarning: Skipping loading of weights for layer conv2d_75 due to mismatch in shape ((1, 1, 256, 156) vs (255, 256, 1, 1)). weight_values[i].shape)) C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\saving.py:1140: UserWarning: Skipping loading of weights for layer conv2d_75 due to mismatch in shape ((156,) vs (255,)). weight_values[i].shape)) Load weights model_data/yolo.h5. Freeze the first 249 layers of total 252 layers. WARNING:tensorflow:From C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\backend\tensorflow_backend.py:1521: The name tf.log is deprecated. Please use tf.math.log instead. WARNING:tensorflow:From C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\backend\tensorflow_backend.py:3080: where (from tensorflow.python.ops.array_ops) is deprecated and will be removed in a future version. Instructions for updating: Use tf.where in 2.0, which has the same broadcast rule as np.where WARNING:tensorflow:From C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\optimizers.py:790: The name tf.train.Optimizer is deprecated. Please use tf.compat.v1.train.Optimizer instead. Train on 1344 samples, val on 335 samples, with batch size 2. WARNING:tensorflow:From C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\backend\tensorflow_backend.py:986: The name tf.assign_add is deprecated. Please use tf.compat.v1.assign_add instead. WARNING:tensorflow:From C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\backend\tensorflow_backend.py:973: The name tf.assign is deprecated. Please use tf.compat.v1.assign instead. WARNING:tensorflow:From C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\callbacks.py:850: The name tf.summary.merge_all is deprecated. Please use tf.compat.v1.summary.merge_all instead. WARNING:tensorflow:From C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\callbacks.py:853: The name tf.summary.FileWriter is deprecated. Please use tf.compat.v1.summary.FileWriter instead. Epoch 1/250 2023-11-19 18:02:57.390440: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:533] shape_optimizer failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1) 2023-11-19 18:02:57.422375: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:533] remapper failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1) 2023-11-19 18:02:57.536783: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:533] layout failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1) 2023-11-19 18:02:57.642200: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:533] shape_optimizer failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1) 2023-11-19 18:02:57.667959: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:533] remapper failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1) 2023-11-19 18:02:57.874195: I tensorflow/stream_executor/platform/default/dso_loader.cc:44] Successfully opened dynamic library cudnn64_7.dll 2023-11-19 18:04:54.112161: W tensorflow/stream_executor/cuda/redzone_allocator.cc:312] Internal: Invoking ptxas not supported on Windows Relying on driver to perform ptx compilation. This message will be only logged once. 2023-11-19 18:04:54.211404: I tensorflow/stream_executor/platform/default/dso_loader.cc:44] Successfully opened dynamic library cublas64_100.dll 2023-11-19 18:05:31.922039: E tensorflow/stream_executor/cuda/cuda_blas.cc:428] failed to run cuBLAS routine: CUBLAS_STATUS_EXECUTION_FAILED Traceback (most recent call last): File "c:/Users/kk/kkFiles/keras-yolo3/train.py", line 206, in <module> _main() File "c:/Users/kk/kkFiles/keras-yolo3/train.py", line 75, in _main callbacks=[logging, checkpoint]) File "C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\legacy\interfaces.py", line 91, in wrapper return func(*args, **kwargs) File "C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\training.py", line 1418, in fit_generator initial_epoch=initial_epoch) File "C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\training_generator.py", line 217, in fit_generator class_weight=class_weight) File "C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\training.py", line 1217, in train_on_batch outputs = self.train_function(ins) File "C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\backend\tensorflow_backend.py", line 2715, in __call__ return self._call(inputs) File "C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\backend\tensorflow_backend.py", line 2675, in _call fetched = self._callable_fn(*array_vals) File "C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\tensorflow_core\python\client\session.py", line 1472, in __call__ run_metadata_ptr) tensorflow.python.framework.errors_impl.InternalError: 2 root error(s) found. (0) Internal: Blas SGEMM launch failed : m=86528, n=32, k=64 [[{{node conv2d_3/convolution}}]] [[loss/add_74/_2891]] (1) Internal: Blas SGEMM launch failed : m=86528, n=32, k=64 [[{{node conv2d_3/convolution}}]] 0 successful operations. 0 derived errors ignored.
已尝试的方法
使用以下代码开启GPU内存动态增长,但未解决问题:
physical_devices = tf.config.experimental.list_physical_devices('GPU') assert len(physical_devices) > 0, "Not enough GPU hardware devices available" tf.config.experimental.set_memory_growth(physical_devices[0], True)
解决思路
- 升级CUDA和cuDNN版本:RTX3070属于安培架构,CUDA10.0不支持该架构,至少需要升级到CUDA11.0及以上版本,搭配对应兼容的cuDNN(比如CUDA11.2配cuDNN8.1);若坚持用TF1.x,最高支持CUDA10.2,但安培架构在TF1.x下兼容性差,建议切换到TF2.x版本。
- 降低batch size:当前batch size为2,可尝试改为1,减少GPU内存占用,避免因内存不足导致cuBLAS调用失败。
- 修复权重加载问题:报错中存在多个层权重形状不匹配的警告,说明预训练权重
yolo.h5是针对80类训练的(255=3*(80+5)),而你的任务是47类(156=3*(47+5)),需要重新生成对应类别的预训练权重,或者修改模型加载逻辑,只加载匹配的层权重。 - 限制GPU内存占用:除了动态增长,还可以尝试固定分配GPU内存,示例代码如下:
from tensorflow.compat.v1 import ConfigProto from tensorflow.compat.v1 import InteractiveSession config = ConfigProto() config.gpu_options.per_process_gpu_memory_fraction = 0.7 # 分配70%的GPU内存 session = InteractiveSession(config=config) - 更新NVIDIA驱动:确保显卡驱动是最新版本,适配你的CUDA版本和RTX3070显卡。
- 检查TensorFlow GPU检测:运行以下代码确认TensorFlow能正确识别GPU:
如果输出中没有GPU设备,说明TensorFlow GPU版本未正确安装,需重新安装对应版本的tensorflow-gpu。from tensorflow.python.client import device_lib print(device_lib.list_local_devices())
内容的提问来源于stack exchange,提问作者user22946469
相关产品推荐
相关产品推荐

