You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

YOLOv3(Keras) GPU训练遇cuBLAS_STATUS_EXECUTION_FAILED错误求助

问题

环境配置

  • 操作系统:Windows 11 Pro
  • GPU:NVIDIA GeForce RTX 3070
  • Python:3.7.16
  • CUDA:10.0.130
  • cuDNN:7.6.5
  • TensorFlow:1.15.0、tensorflow-gpu 1.15.0
  • Keras:2.2.4

报错情况

运行YOLOv3训练脚本时无法使用GPU训练,报错failed to run cuBLAS routine: CUBLAS_STATUS_EXECUTION_FAILED,完整报错输出如下:

Create YOLOv3 model with 9 anchors and 47 classes.
C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\saving.py:1140: UserWarning: Skipping loading of weights for layer conv2d_59 due to mismatch in shape ((1, 1, 1024, 156) vs (255, 1024, 1, 1)).
weight_values[i].shape))
C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\saving.py:1140: UserWarning: Skipping loading of weights for layer conv2d_59 due to mismatch in shape ((156,) vs (255,)).
weight_values[i].shape))
C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\saving.py:1140: UserWarning: Skipping loading of weights for layer conv2d_67 due to mismatch in shape ((1, 1, 512, 156) vs (255, 512, 1, 1)).
weight_values[i].shape))
C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\saving.py:1140: UserWarning: Skipping loading of weights for layer conv2d_67 due to mismatch in shape ((156,) vs (255,)).
weight_values[i].shape))
C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\saving.py:1140: UserWarning: Skipping loading of weights for layer conv2d_75 due to mismatch in shape ((1, 1, 256, 156) vs (255, 256, 1, 1)).
weight_values[i].shape))
C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\saving.py:1140: UserWarning: Skipping loading of weights for layer conv2d_75 due to mismatch in shape ((156,) vs (255,)).
weight_values[i].shape))
Load weights model_data/yolo.h5.
Freeze the first 249 layers of total 252 layers.
WARNING:tensorflow:From C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\backend\tensorflow_backend.py:1521: The name tf.log is deprecated. Please use tf.math.log instead.

WARNING:tensorflow:From C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\backend\tensorflow_backend.py:3080: where (from tensorflow.python.ops.array_ops) is deprecated and will be removed in a future version.
Instructions for updating:
Use tf.where in 2.0, which has the same broadcast rule as np.where
WARNING:tensorflow:From C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\optimizers.py:790: The name tf.train.Optimizer is deprecated. Please use tf.compat.v1.train.Optimizer instead.

Train on 1344 samples, val on 335 samples, with batch size 2.
WARNING:tensorflow:From C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\backend\tensorflow_backend.py:986: The name tf.assign_add is deprecated. Please use tf.compat.v1.assign_add instead.

WARNING:tensorflow:From C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\backend\tensorflow_backend.py:973: The name tf.assign is deprecated. Please use tf.compat.v1.assign instead.

WARNING:tensorflow:From C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\callbacks.py:850: The name tf.summary.merge_all is deprecated. Please use tf.compat.v1.summary.merge_all instead.

WARNING:tensorflow:From C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\callbacks.py:853: The name tf.summary.FileWriter is deprecated. Please use tf.compat.v1.summary.FileWriter instead.

Epoch 1/250
2023-11-19 18:02:57.390440: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:533] shape_optimizer failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1)
2023-11-19 18:02:57.422375: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:533] remapper failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1)
2023-11-19 18:02:57.536783: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:533] layout failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1)
2023-11-19 18:02:57.642200: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:533] shape_optimizer failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1)
2023-11-19 18:02:57.667959: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:533] remapper failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1)
2023-11-19 18:02:57.874195: I tensorflow/stream_executor/platform/default/dso_loader.cc:44] Successfully opened dynamic library cudnn64_7.dll
2023-11-19 18:04:54.112161: W tensorflow/stream_executor/cuda/redzone_allocator.cc:312] Internal: Invoking ptxas not supported on Windows
Relying on driver to perform ptx compilation. This message will be only logged once.
2023-11-19 18:04:54.211404: I tensorflow/stream_executor/platform/default/dso_loader.cc:44] Successfully opened dynamic library cublas64_100.dll
2023-11-19 18:05:31.922039: E tensorflow/stream_executor/cuda/cuda_blas.cc:428] failed to run cuBLAS routine: CUBLAS_STATUS_EXECUTION_FAILED
Traceback (most recent call last):
File "c:/Users/kk/kkFiles/keras-yolo3/train.py", line 206, in <module>
_main()
File "c:/Users/kk/kkFiles/keras-yolo3/train.py", line 75, in _main
callbacks=[logging, checkpoint])
File "C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\legacy\interfaces.py", line 91, in wrapper
return func(*args, **kwargs)
File "C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\training.py", line 1418, in fit_generator
initial_epoch=initial_epoch)
File "C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\training_generator.py", line 217, in fit_generator
class_weight=class_weight)
File "C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\engine\training.py", line 1217, in train_on_batch
outputs = self.train_function(ins)
File "C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\backend\tensorflow_backend.py", line 2715, in __call__
return self._call(inputs)
File "C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\keras\backend\tensorflow_backend.py", line 2675, in _call
fetched = self._callable_fn(*array_vals)
File "C:\Users\kk\.conda\envs\kk_Environment\lib\site-packages\tensorflow_core\python\client\session.py", line 1472, in __call__
run_metadata_ptr)
tensorflow.python.framework.errors_impl.InternalError: 2 root error(s) found.
(0) Internal: Blas SGEMM launch failed : m=86528, n=32, k=64
[[{{node conv2d_3/convolution}}]]
[[loss/add_74/_2891]]
(1) Internal: Blas SGEMM launch failed : m=86528, n=32, k=64
[[{{node conv2d_3/convolution}}]]
0 successful operations.
0 derived errors ignored.

已尝试的方法

使用以下代码开启GPU内存动态增长,但未解决问题:

physical_devices = tf.config.experimental.list_physical_devices('GPU')
assert len(physical_devices) > 0, "Not enough GPU hardware devices available"
tf.config.experimental.set_memory_growth(physical_devices[0], True)

解决思路

  • 升级CUDA和cuDNN版本:RTX3070属于安培架构,CUDA10.0不支持该架构,至少需要升级到CUDA11.0及以上版本,搭配对应兼容的cuDNN(比如CUDA11.2配cuDNN8.1);若坚持用TF1.x,最高支持CUDA10.2,但安培架构在TF1.x下兼容性差,建议切换到TF2.x版本。
  • 降低batch size:当前batch size为2,可尝试改为1,减少GPU内存占用,避免因内存不足导致cuBLAS调用失败。
  • 修复权重加载问题:报错中存在多个层权重形状不匹配的警告,说明预训练权重yolo.h5是针对80类训练的(255=3*(80+5)),而你的任务是47类(156=3*(47+5)),需要重新生成对应类别的预训练权重,或者修改模型加载逻辑,只加载匹配的层权重。
  • 限制GPU内存占用:除了动态增长,还可以尝试固定分配GPU内存,示例代码如下:
    from tensorflow.compat.v1 import ConfigProto
    from tensorflow.compat.v1 import InteractiveSession
    
    config = ConfigProto()
    config.gpu_options.per_process_gpu_memory_fraction = 0.7  # 分配70%的GPU内存
    session = InteractiveSession(config=config)
    
  • 更新NVIDIA驱动:确保显卡驱动是最新版本,适配你的CUDA版本和RTX3070显卡。
  • 检查TensorFlow GPU检测:运行以下代码确认TensorFlow能正确识别GPU:
    from tensorflow.python.client import device_lib
    print(device_lib.list_local_devices())
    
    如果输出中没有GPU设备,说明TensorFlow GPU版本未正确安装,需重新安装对应版本的tensorflow-gpu。

内容的提问来源于stack exchange,提问作者user22946469

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 14:10:54