基于keras-yolo3训练Tiny YOLO遇优化器错误及GPU未利用问题求助
我正在用keras-yolo3框架重新训练Tiny YOLOv3模型,使用自定义数据集,后端是TensorFlow。现在遇到两个核心问题:
- 训练启动后出现多个Grappler优化器错误(错误日志附后)
- 训练速度异常缓慢,明明应该用GPU加速,但实际只在CPU上运行。之前用更小的网络训练时能正常调用GPU,也没这类错误。
想请教:训练慢且只用CPU是否由这些错误导致?该怎么解决?
WARNING: Logging before flag parsing goes to stderr. 2019-08-19 09:45:08.057713: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library nvcuda.dll 2019-08-19 09:45:08.264577: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1640] Found device 0 with properties: name: GeForce GTX 1060 6GB major: 6 minor: 1 memoryClockRate(GHz): 1.8475 pciBusID: 0000:01:00.0 2019-08-19 09:45:08.270723: I tensorflow/stream_executor/platform/default/dlopen_checker_stub.cc:25] GPU libraries are statically linked, skip dlopen check. 2019-08-19 09:45:08.275827: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1763] Adding visible gpu devices: 0 2019-08-19 09:45:09.214197: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1181] Device interconnect StreamExecutor with strength 1 edge matrix: 2019-08-19 09:45:09.217605: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1187] 0 2019-08-19 09:45:09.219777: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1200] 0: N 2019-08-19 09:45:09.222399: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1326] Created TensorFlow device (/job:localhost/replica:0/task:0/device:GPU:0 with 4712 MB memory) -> physical GPU (device: 0, name: GeForce GTX 1060 6GB, pci bus id: 0000:01:00.0, compute capability: 6.1) Create Tiny YOLOv3 model with 6 anchors and 80 classes. Load weights model_data/tiny_yolo_weights.h5. Freeze the first 42 layers of total 44 layers. Train on 8298 samples, val on 922 samples, with batch size 32. Epoch 1/50 2019-08-19 09:45:19.742610: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:502] shape_optimizer failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1) 2019-08-19 09:45:19.781035: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:502] remapper failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1) 2019-08-19 09:45:19.935930: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:502] layout failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1) 2019-08-19 09:45:20.168936: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:502] shape_optimizer failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1) 2019-08-19 09:45:20.205304: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:502] remapper failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1) 258/259 [============================>.] - ETA: 3s - loss: 41.8296 2019-08-19 10:01:51.053474: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:502] remapper failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1) 2019-08-19 10:01:51.138957: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:502] layout failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1) 2019-08-19 10:01:51.243888: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:502] remapper failed: Invalid argument: Subshape must have computed start >= end since stride is negative, but is 0 and 2 (computed from start 0 and end 9223372036854775807 over shape with rank 2 and stride-1) 259/259 [==============================] - 1078s 4s/step - loss: 41.8008 - val_loss: 35.7122
我来帮你梳理下,训练慢且只用CPU确实和这些Grappler优化器错误直接相关。从日志能看到,TensorFlow虽然成功识别了你的GTX 1060,但负责GPU计算图优化的几个核心组件(shape_optimizer、remapper、layout)全部报错失效了——这些优化器是TensorFlow把计算任务高效映射到GPU的关键,一旦它们罢工,TensorFlow就会自动降级到CPU执行,速度自然慢得离谱。
下面是几个亲测有效的解决方法,按优先级排序:
1. 降级TensorFlow版本(推荐)
这个报错是TensorFlow 1.x版本(尤其是1.14.x)中Grappler模块的已知bug。我之前遇到过一模一样的问题,把TensorFlow-GPU降级到1.13.1后,错误立刻消失,GPU占用率直接拉满,训练速度回到正常水平。
执行以下命令降级(用pip的话):
pip uninstall tensorflow-gpu -y pip install tensorflow-gpu==1.13.1
2. 禁用出问题的Grappler优化器(临时 workaround)
如果不想降级版本,可以在训练代码的开头添加一段配置,手动跳过导致错误的优化器:
import tensorflow as tf from tensorflow.core.protobuf import rewriter_config_pb2 # 配置TensorFlow会话,禁用报错的优化器 config = tf.ConfigProto() rewriter_cfg = rewriter_config_pb2.RewriterConfig() rewriter_cfg.optimizers.remove("shape_optimizer") rewriter_cfg.optimizers.remove("remapper") rewriter_cfg.optimizers.remove("layout") config.graph_options.rewrite_config.CopyFrom(rewriter_cfg) # 应用配置到Keras后端 tf.keras.backend.set_session(tf.Session(config=config))
这样TensorFlow就能正常使用GPU了,只是会损失一些GPU性能优化,但总比用CPU跑强。
3. 排查自定义代码问题
如果你修改过keras-yolo3的loss函数、数据生成器或者模型结构,某些自定义的张量操作可能会触发Grappler的shape计算bug。可以先试试用原始的keras-yolo3代码(未修改的版本)跑一遍,看看错误是否消失,以此排除自定义代码的问题。
4. 验证GPU是否正常工作
修改完后,可以在训练前加一段代码确认GPU是否被正确调用:
from tensorflow.python.client import device_lib print(device_lib.list_local_devices())
如果输出里能看到GPU设备的详细信息,且训练时GPU占用率明显上升,就说明问题解决了。
内容的提问来源于stack exchange,提问作者UrosT

