YOLOv3自定义目标训练原理及动漫人物检测性能优化咨询
YOLOv3 Custom Training & Speed Optimization for Anime Character Detection
1. How YOLOv3 Custom Object Training Works
Let's break down the core workflow in plain terms, since you're working on an anime character detection project:
- Dataset Preparation:
- Label your anime frames with tools like LabelImg, generating a
.txtannotation file for every image. Each line in the txt follows this format:class_index x_center y_center width height(all coordinates are normalized to 0-1, so they work with any image size). - Create
train.txt,val.txtfiles that list the absolute paths of your training/validation images—this tells the model which data to learn from.
- Label your anime frames with tools like LabelImg, generating a
- Config File Tweaks:
- Copy the original
yolov3.cfgand modify key settings for your use case:- Set
classesto 1 (since you only need to detect characters). - Adjust
filtersin each YOLO layer to(classes + 5) * 3(that's 18 for 1 class—this formula accounts for bounding box coordinates, confidence, and class probabilities). - Update
train,valid,names, andbackuppaths to point to your dataset files and where you want to save trained weights. - For GTX 1050Ti, set
batch=16andsubdivisions=8to avoid out-of-memory errors.
- Set
- Copy the original
- Pretrained Weight Initialization:
- Use
darknet53.conv.74(the pre-trained backbone weights) instead of the fullyolov3.weights. This lets the model reuse general image features (like edges, shapes) learned from large datasets, so it trains faster on your anime data.
- Use
- Training Process:
- Run this command in the Darknet framework:
./darknet detector train data/your_anime_data.data cfg/your_custom_yolov3.cfg darknet53.conv.74 - Watch the loss value—stop training once it stabilizes (usually around 0.1-0.5 after a few thousand iterations). The trained weights will save to your
backupfolder automatically.
- Run this command in the Darknet framework:
- Inference & Validation:
- Test your model with:
./darknet detector test data/your_anime_data.data cfg/your_custom_yolov3.cfg backup/your_custom_yolov3_final.weights test_frame.jpg - Tweak the
confidencethreshold in the cfg file to filter out low-confidence, false detections.
- Test your model with:
2. Speed Optimization for GTX 1050Ti (Target: 10 FPS)
First off, running full YOLOv3 at 608×608 on a GTX 1050Ti is always going to be slow—1.5 FPS is totally expected. Let's tackle your two questions and add actionable tips to hit that 10 FPS mark:
① Does the number of classes affect detection speed?
Yes, but nowhere near an 80x speedup. Here's the real story:
- ~85% of inference time is spent in the Darknet53 backbone (the part that extracts image features), and this doesn't care how many classes you're detecting.
- The only speed gain comes from the YOLO detection layers: each anchor box calculates
classes + 5values (bounding box coords, confidence, class probabilities). Cutting classes from 80 to 1 reduces this computation, but it only gives a 10-20% speed boost at most. - That said, modifying the cfg to set
classes=1andfilters=18is still a quick, no-downside win for your project.
② Should I adjust the original 1920×1080 training input resolution?
Absolutely—this is one of the most impactful changes you can make, and it's required for proper YOLOv3 training:
- YOLOv3 requires input resolutions to be multiples of 32 (thanks to 5 rounds of 2x downsampling in the backbone, totaling 32x downsampling). 1080 isn't a multiple of 32, so resizing directly would cause alignment issues in the network.
- For speed, drop the resolution to a smaller multiple of 32 like 416×416 or 320×320. YOLOv3's compute time scales with the square of the input resolution:
- Going from 608×608 to 416×416 cuts compute to ~47% of the original, which should boost FPS to ~3-4x (around 5-7 FPS on your GPU).
- Anime characters have distinct, simple features, so you likely won't lose much detection accuracy with smaller resolutions—test 416 first, then 320 if you need more speed.
- When training, resize your 1920×1080 images to your target resolution (or use random cropping + resizing as data augmentation) to match the inference input.
Bonus Tips to Hit 10 FPS
- Switch to YOLOv3-Tiny: This lightweight variant strips out most Darknet53 layers and uses fewer anchors. On GTX 1050Ti, it can hit 15-20 FPS at 416×416, with only a minor accuracy drop for anime characters.
- Use TensorRT Optimization: Convert your YOLOv3 model to TensorRT format (FP16 precision) to leverage GPU-specific acceleration. This can double your FPS without losing meaningful accuracy.
- Optimize Inference Settings: Set
batch=1during inference, disable verbose logging, and use thethreshparameter to skip low-confidence detections early.
内容的提问来源于stack exchange,提问作者Akinohana
相关产品推荐
相关产品推荐

