TensorFlow目标检测训练启动即崩溃问题求助
我尝试使用自定义数据集训练图像检测器,但训练无法启动。除类别数量和路径外,我对配置文件做了如下修改:
train_input_reader: { tf_record_input_reader { input_path: "data/data_train.record" } label_map_path: "data/rdata_train.pbtxt" queue_capacity: 2 min_after_dequeue: 1 } train_config: { batch_size: 1 optimizer { rms_prop_optimizer: { learning_rate: { exponential_decay_learning_rate { initial_learning_rate: 0.1 decay_steps: 800720 decay_factor: 0.95 } } momentum_optimizer_value: 0.9 decay: 0.9 epsilon: 0.000001 } } fine_tune_checkpoint: "ssd_mobilenet_v1_coco_2017_11_17/model.ckpt" from_detection_checkpoint: true # Note: The below line limits the training process to 200K steps, which we # empirically found to be sufficient enough to train the pets dataset. This # effectively bypasses the learning rate schedule (the learning rate will # never decay). Remove the below line to train indefinitely. num_steps: 200000 batch_queue_capacity: 5 num_batch_queue_threads: 8 prefetch_queue_capacity: 5 }
但训练始终无法启动,我已查阅TensorFlow Research仓库的相关问题,但未找到针对性解决方案。控制台输出的开头内容如下:
/home/ubuntu/.conda/envs/tf/lib/python3.6/site-packages/h5py/init.py:36: FutureWarning: Conversion of the second argument of issubdtype from
floattonp.floatingis deprecated. In future, it will be treated asnp.float64 == np.dtype(float).type.
from ._conv impo...
问题排查与解决方案
首先得说,那个h5py的警告只是版本兼容提示,不是训练启动失败的直接原因,不用太纠结它。咱们重点看配置和数据环节的几个潜在问题:
队列参数设置过小:你把
queue_capacity设为2、min_after_dequeue设为1,batch_queue_capacity和prefetch_queue_capacity也只有5,这种极小的队列容量很容易导致数据读取阻塞——线程还没加载到足够数据,训练进程就卡住启动不了。建议调整这些参数:queue_capacity改成1000min_after_dequeue改成500batch_queue_capacity改成50prefetch_queue_capacity改成20
这样能保证数据队列有足够缓存,支撑训练的持续数据读取。
学习率设置过高:SSD MobileNet这类轻量模型的初始学习率一般用0.001就足够了,你设成0.1实在太高,会直接导致模型参数震荡,甚至一开始就出现NaN值让训练崩溃。赶紧把
initial_learning_rate改成0.001,同时decay_steps也要对应调整(比如总步数20万的话,设成20000,也就是总步数的10%左右),让学习率衰减更合理。验证数据集文件有效性:虽然你确认了路径,但还是要检查
data_train.record和rdata_train.pbtxt是否正常:- 用TensorFlow的工具检查record文件是否损坏,能否正常读取样本
- 确认label map里的类别ID是连续的,且和训练标注完全对应
线程数与队列容量匹配:你设置了
num_batch_queue_threads: 8,但队列容量太小,线程太多反而会引发资源竞争。建议先把线程数降到4,等队列参数调大后再根据实际情况调整。
如果调整后还是无法启动,把控制台的完整报错信息贴出来,这样能更精准定位问题。
内容的提问来源于stack exchange,提问作者Aditya Thakkar

