You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取车牌每个字符的Bounding Box?用于训练字符检测神经网络

Solutions for Generating License Plate Character Bounding Boxes (PASCAL VOC XML)

Hey Vlad, sorry to hear you've hit roadblocks with traditional CV methods and initial detection network attempts. Let's walk through practical, robust approaches to get those character bounding boxes for your 10k license plate dataset.

Since you have known character labels but no BBoxes, small-scale manual annotation + model fine-tuning is the most reliable way to handle the lighting/noise variations in your dataset. Here's how to do it with YOLOv8 (easy to implement, fast to train):

  • Step 1: Manual Annotation (Small Sample)
    Grab 200-300 representative images (cover different lighting/contrast scenarios) and use LabelImg to draw BBoxes for each character. Save annotations in PASCAL VOC format directly, or convert to YOLO format later.

  • Step 2: Fine-Tune YOLOv8
    Use the YOLOv8n (nano) model for speed, since your task is focused on small, uniform objects (license plate characters):

    from ultralytics import YOLO
    
    # Load pre-trained model
    model = YOLO('yolov8n.pt')
    
    # Fine-tune on your annotated dataset
    model.train(data='your_dataset.yaml', epochs=50, imgsz=640, batch=16)
    

    Your dataset.yaml should define train/val splits, class names (all possible license plate characters: А, В, Е, К, М, Н, О, Р, С, Т, У, Х, 0-9), and data paths.

  • Step 3: Batch Predict & Convert to PASCAL VOC XML
    Run the fine-tuned model on your remaining 9700+ images to get BBox predictions. Then write a simple script to convert YOLO's output (class + x1,y1,x2,y2) into PASCAL VOC XML format, mapping each detected character to your known text labels (match by position order, since characters are left-to-right).

2. Adaptive Preprocessing + Rule-Based Contour Detection (No Annotation Needed)

Leverage the fixed structure of license plates (characters are horizontally aligned, uniform size range) to bypass full model training:

  • Step 1: Contrast Enhancement & Adaptive Thresholding
    Fix lighting/contrast issues first with CLAHE (adaptive histogram equalization), then use adaptive thresholding instead of global binary segmentation:

    import cv2
    
    img = cv2.imread('В394ТТ64.png', cv2.IMREAD_GRAYSCALE)
    # Enhance contrast
    clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8,8))
    img_enhanced = clahe.apply(img)
    # Adaptive thresholding
    thresh = cv2.adaptiveThreshold(img_enhanced, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)
    
  • Step 2: Morphological Cleaning & Contour Filtering
    Remove noise with morphological operations, then extract and filter contours based on size and aspect ratio (tailor these values to your license plate character dimensions):

    # Remove small noise
    kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (2,2))
    cleaned = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, kernel)
    # Find contours
    contours, _ = cv2.findContours(cleaned, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
    # Filter contours: keep those with width/height in expected range, aspect ratio ~1:2 (adjust based on your plates)
    valid_contours = []
    for cnt in contours:
        x, y, w, h = cv2.boundingRect(cnt)
        if 20 < w < 60 and 40 < h < 100 and 0.4 < w/h < 0.6:
            valid_contours.append((x, y, w, h))
    
  • Step 3: Sort Contours & Map to Known Characters
    Sort valid contours by their x-coordinate (left to right) to match the order of your known character text. Each contour's bounding box is your character BBox.

3. Tesseract OCR with Custom Character Whitelist

If you don't want to train a model or do manual annotation, use Tesseract with targeted preprocessing and a character whitelist to extract BBoxes:

  • Step 1: Configure Tesseract for License Plates
    Install Tesseract with Russian language support, then set a whitelist of allowed license plate characters to reduce false detections:

    import pytesseract
    
    custom_config = r'--oem 3 --psm 8 -c tessedit_char_whitelist=АВЕКМНОРСТУХ0123456789'
    # --psm 8 treats the image as a single word, adjust to 7 if needed
    
  • Step 2: Extract BBox Data
    Use image_to_data to get detailed character-level BBox info:

    # Use the enhanced/thresholded image from Method 2
    results = pytesseract.image_to_data(img_enhanced, config=custom_config, output_type=pytesseract.Output.DICT)
    # Extract BBoxes for valid characters (confidence > 50%)
    char_bboxes = []
    for i in range(len(results['text'])):
        if int(results['conf'][i]) > 50 and results['text'][i].strip() != '':
            x = results['left'][i]
            y = results['top'][i]
            w = results['width'][i]
            h = results['height'][i]
            char_bboxes.append((x, y, w, h, results['text'][i]))
    
  • Step 3: Align with Known Text
    Sort the detected characters by x-coordinate, then match them to your known label text. If there's a mismatch, adjust preprocessing or Tesseract parameters (e.g., psm mode, whitelist).

Final Tips

  • Post-Validation: After generating BBoxes, write a script to check if the number of BBoxes matches the length of your known character text, and that BBoxes are horizontally aligned. Fix any outliers manually (or use a small correction script).
  • Hybrid Approach: If Method 2 has edge cases, use it to generate initial BBoxes, then manually correct 100-200 bad ones, and fine-tune a model on that combined dataset for better accuracy.

内容的提问来源于stack exchange,提问作者Vlad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 10:07:58