如何在TensorFlow 2.0中正确修改目标检测模型的训练epoch数量
项目基础配置
我正在搭建仅用于车辆检测的模型,使用配置如下:
- 技术方案:搭载迁移学习的TensorFlow 2.0
- 预训练模型:ssd_mobilenet_v2_fpnlite_320x320_coco17_tpu-8
- 数据集:仅保留car类别的coco-2017数据集
训练现象
训练完成后运行模型,准确率约为50%,total_loss呈下降趋势但未收敛,可查看检测车辆结果和tensorboard记录作为参考。
total_loss下降是正向表现,但模型进一步收敛可获得更高准确率。我尝试将pipeline.config中的num_epochs参数从1修改为4后没有生效,推测该参数不是调整训练轮次的正确配置项。
核心疑问
如何正确修改TensorFlow 2.0下模型训练的epoch数量?
现有pipeline.config配置
model { ssd { num_classes: 1 image_resizer { fixed_shape_resizer { height: 320 width: 320 } } feature_extractor { type: "ssd_mobilenet_v2_fpn_keras" depth_multiplier: 1.0 min_depth: 16 conv_hyperparams { regularizer { l2_regularizer { weight: 4e-05 } } initializer { random_normal_initializer { mean: 0.0 stddev: 0.01 } } activation: RELU_6 batch_norm { decay: 0.997 scale: true epsilon: 0.001 } } use_depthwise: true override_base_feature_extractor_hyperparams: true fpn { min_level: 3 max_level: 7 additional_layer_depth: 128 } } box_coder { faster_rcnn_box_coder { y_scale: 10.0 x_scale: 10.0 height_scale: 5.0 width_scale: 5.0 } } matcher { argmax_matcher { matched_threshold: 0.5 unmatched_threshold: 0.5 ignore_thresholds: false negatives_lower_than_unmatched: true force_match_for_each_row: true use_matmul_gather: true } } similarity_calculator { iou_similarity { } } box_predictor { weight_shared_convolutional_box_predictor { conv_hyperparams { regularizer { l2_regularizer { weight: 4e-05 } } initializer { random_normal_initializer { mean: 0.0 stddev: 0.01 } } activation: RELU_6 batch_norm { decay: 0.997 scale: true epsilon: 0.001 } } depth: 128 num_layers_before_predictor: 4 kernel_size: 3 class_prediction_bias_init: -4.6 share_prediction_tower: true use_depthwise: true } } anchor_generator { multiscale_anchor_generator { min_level: 3 max_level: 7 anchor_scale: 4.0 aspect_ratios: 1.0 aspect_ratios: 2.0 aspect_ratios: 0.5 scales_per_octave: 2 } } post_processing { batch_non_max_suppression { score_threshold: 1e-08 iou_threshold: 0.6 max_detections_per_class: 100 max_total_detections: 100 use_static_shapes: false } score_converter: SIGMOID } normalize_loss_by_num_matches: true loss { localization_loss { weighted_smooth_l1 { } } classification_loss { weighted_sigmoid_focal { gamma: 2.0 alpha: 0.25 } } classification_weight: 1.0 localization_weight: 1.0 } encode_background_as_zeros: true normalize_loc_loss_by_codesize: true inplace_batchnorm_update: true freeze_batchnorm: false } } train_config { batch_size: 4 data_augmentation_options { random_horizontal_flip { } } data_augmentation_options { random_crop_image { min_object_covered: 0.0 min_aspect_ratio: 0.75 max_aspect_ratio: 3.0 min_area: 0.75 max_area: 1.0 overlap_thresh: 0.0 } } sync_replicas: true optimizer { momentum_optimizer { learning_rate { cosine_decay_learning_rate { learning_rate_base: 0.0008 total_steps: 50000 warmup_learning_rate: 0.00026666 warmup_steps: 1000 } } momentum_optimizer_value: 0.9 } use_moving_average: false } fine_tune_checkpoint: "Tensorflow/workspace/pre-trained-models/ssd_mobilenet_v2_fpnlite_320x320_coco17_tpu-8/checkpoint/ckpt-0" num_steps: 50000 startup_delay_steps: 0.0 replicas_to_aggregate: 8 max_number_of_boxes: 100 unpad_groundtruth_tensors: false fine_tune_checkpoint_type: "detection" fine_tune_checkpoint_version: V2 } train_input_reader { label_map_path: "Tensorflow/workspace/annotations/label_map.pbtxt" tf_record_input_reader { input_path: "Tensorflow/workspace/annotations/train.record" } } eval_config { metrics_set: "coco_detection_metrics" use_moving_averages: false } eval_input_reader { label_map_path: "Tensorflow/workspace/annotations/label_map.pbtxt" shuffle: false num_epochs: 4 tf_record_input_reader { input_path: "Tensorflow/workspace/annotations/test.record" } }
你之前修改的num_epochs位于eval_input_reader配置块下,该参数仅控制评估阶段测试集的遍历次数,和训练轮次无关。
TensorFlow 2.x Object Detection API默认不直接配置训练epoch数,而是通过train_config下的num_steps参数换算得到训练轮次,换算公式为:训练轮数 = 总训练步数(num_steps) ÷ 单轮训练步数
其中单轮训练步数 = 训练集总样本数 ÷ 训练batch_size
举个实际计算示例:假设你的训练集共包含12000张标注好的车辆图片,当前配置的batch_size为4,那么单轮训练需要走12000/4=3000步;如果需要训练4轮,直接把train_config下的num_steps参数调整为3000*4=12000即可。
如果想直接指定epoch数启动训练,无需手动换算步数,也可以在启动训练的脚本(通常是model_main_tf2.py)的运行命令中添加--num_train_epochs=4参数,该参数优先级高于pipeline.config中的num_steps配置。
另外针对你当前loss未收敛的情况,除了增加训练轮次,也可以搭配调整学习率base值、补充随机缩放/色域扰动等数据增强策略,进一步提升模型准确率。
内容的提问来源于stack exchange,提问作者Hoang97

