You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于SSD MobileDet与TensorFlow 1.15.5的分类损失优化问询

SSDLite MobileDet自定义模型训练优化问题

我使用TensorFlow 1.15.5,基于Google Coral的SSDLite MobileDet教程,用约20000张标注图像从零训练自定义模型。

我的配置基于官方示例,并做了以下修改:

  • 类别数:11
  • 批次大小:64
  • 基础学习率:0.2
  • 根据训练集自定义宽高比

使用这些参数,我得到了相当不错的结果:
训练结果曲线1
训练结果曲线2
训练结果曲线3
训练结果曲线4
训练结果曲线5

我尝试进一步优化结果,降低呈轻微上升趋势的分类损失。

通常针对损失上升的建议是降低学习率,因此我首先调整学习率如下:

cosine_decay_learning_rate {
  learning_rate_base: 0.12      # 此前为0.2
  total_steps: 400000
  warmup_learning_rate: 0.01333 # 此前为0.1333
  warmup_steps: 281             # ~1轮训练,此前为2000
}

这使得mAP、AR和总损失有所改善,但分类损失在约120000步后开始急剧上升:
分类损失上升曲线1

因此我尝试进一步降低学习率,以期解决该问题:

cosine_decay_learning_rate {
  learning_rate_base: 0.03
  total_steps: 400000
  warmup_learning_rate: 0.00333
  warmup_steps: 281 # 1轮训练
}

结果mAP、AR下降,且分类损失(橙色曲线)更高:
分类损失上升曲线2

这让我认为或许只需将总步数减少至200000,让学习率下降更快,而非初始值更低:

cosine_decay_learning_rate {
  learning_rate_base: 0.12
  total_steps: 200000
  warmup_learning_rate: 0.01333
  warmup_steps: 281 # 1轮训练
}

结果mAP、AR与之前相近,但出乎意料的是,随着学习率下降加快,分类损失上升得更快:
分类损失上升曲线3
分类损失上升曲线4

进一步降低总步数会加速分类损失的上升,我不确定下一步该尝试什么。

以下是我的完整pipeline.config:

model {
  ssd {
    num_classes: 11
    image_resizer {
      fixed_shape_resizer {
        height: 320
        width: 320
      }
    }
    feature_extractor {
      type: "ssd_mobiledet_edgetpu"
      depth_multiplier: 1.0
      min_depth: 16
      conv_hyperparams {
        regularizer {
          l2_regularizer {
            weight: 4e-05
          }
        }
        initializer {
          truncated_normal_initializer {
            mean: 0.0
            stddev: 0.03
          }
        }
        activation: RELU_6
        batch_norm {
          decay: 0.97
          center: true
          scale: true
          epsilon: 0.001
          train: true
        }
      }
      use_depthwise: true
      override_base_feature_extractor_hyperparams: false
    }
    box_coder {
      faster_rcnn_box_coder {
        y_scale: 10.0
        x_scale: 10.0
        height_scale: 5.0
        width_scale: 5.0
      }
    }
    matcher {
      argmax_matcher {
        matched_threshold: 0.5
        unmatched_threshold: 0.5
        ignore_thresholds: false
        negatives_lower_than_unmatched: true
        force_match_for_each_row: true
        use_matmul_gather: true
      }
    }
    similarity_calculator {
      iou_similarity {
      }
    }
    box_predictor {
      convolutional_box_predictor {
        conv_hyperparams {
          regularizer {
            l2_regularizer {
              weight: 4e-05
            }
          }
          initializer {
            random_normal_initializer {
              mean: 0.0
              stddev: 0.03
            }
          }
          activation: RELU_6
          batch_norm {
            decay: 0.97
            center: true
            scale: true
            epsilon: 0.001
            train: true
          }
        }
        min_depth: 0
        max_depth: 0
        num_layers_before_predictor: 0
        use_dropout: false
        dropout_keep_probability: 0.8
        kernel_size: 3
        box_code_size: 4
        apply_sigmoid_to_scores: false
        class_prediction_bias_init: -4.6
        use_depthwise: true
      }
    }
    anchor_generator {
      ssd_anchor_generator {
        num_layers: 6
        min_scale: 0.0625
        max_scale: 0.95
        aspect_ratios: 0.44
        aspect_ratios: 0.85
        aspect_ratios: 1.5
        aspect_ratios: 2.41
      }
    }
    post_processing {
      batch_non_max_suppression {
        score_threshold: 1e-08
        iou_threshold: 0.6
        max_detections_per_class: 100
        max_total_detections: 100
        use_static_shapes: true
      }
      score_converter: SIGMOID
    }
    normalize_loss_by_num_matches: true
    loss {
      localization_loss {
        weighted_smooth_l1 {
          delta: 1.0
        }
      }
      classification_loss {
        weighted_sigmoid_focal {
          gamma: 2.0
          alpha: 0.75
        }
      }
      classification_weight: 1.0
      localization_weight: 1.0
    }
    encode_background_as_zeros: true
    normalize_loc_loss_by_codesize: true
    inplace_batchnorm_update: true
    freeze_batchnorm: false
  }
}
train_config {
  batch_size: 64
  data_augmentation_options {
    random_horizontal_flip { 
    }
  }
  data_augmentation_options {
    ssd_random_crop_fixed_aspect_ratio {
    }
  }
  sync_replicas: true
  optimizer {
    momentum_optimizer {
      learning_rate {
        cosine_decay_learning_rate {
          learning_rate_base: 0.12
          total_steps: 400000
          warmup_learning_rate: 0.01333
          warmup_steps: 281 # 1 epoch
        }
      }
      momentum_optimizer_value: 0.9
    }
    use_moving_average: false
  }
  num_steps: 400000
  startup_delay_steps: 0.0
  replicas_to_aggregate: 32
  max_number_of_boxes: 100
  unpad_groundtruth_tensors: false
}
train_input_reader {
  label_map_path: "/lab/labelmap.pbtxt"
  tf_record_input_reader {
    input_path: "/lab/data/train.records"
  }
}
eval_config {
  num_examples: 8000
  metrics_set: "coco_detection_metrics"
  use_moving_averages: false
}
eval_input_reader {
  label_map_path: "/lab/labelmap.pbtxt"
  shuffle: false
  num_epochs: 1
  tf_record_input_reader {
    input_path: "/lab/data/val.records"
  }
}
graph_rewriter {
  quantization {
    delay: 0
    weight_bits: 8
    activation_bits: 8
  }
}

内容的提问来源于stack exchange,提问作者Blake B

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 18:40:26