You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

乌尔都语Tesseract 4.00训练流程及图像单词标注方法咨询

Great to hear you've got Tesseract 4.0 set up and a solid Urdu dataset with those popular Nastaleeq fonts—those are tricky but super useful for training a robust model! Let's break down your two questions step by step:

1. Complete Tesseract Training Workflow for Urdu

Tesseract 4.0 uses an LSTM-based model, so we'll focus on fine-tuning (or building from scratch) a model optimized for your Nastaleeq fonts. Here's the full process tailored to your setup:

Step 1: Preprocess Your Dataset

  • Image Normalization: Ensure all images have consistent resolution (e.g., 300 DPI), clean up noise (using tools like OpenCV's bilateralFilter), and confirm text is correctly oriented (Urdu is right-to-left, so no upside-down or rotated images).
  • Clean Annotations: Double-check your labeled text to remove extra spaces, typos, or non-Urdu characters. For Nastaleeq fonts, pay close attention to ligatures (joined characters) — your annotations should match the exact rendered text (e.g., بعد نجی ٹی وی instead of incorrectly splitting connected characters).
  • Split Data: Divide your dataset into training (80-90%) and evaluation (10-20%) sets to monitor model performance during training.

Step 2: Generate Tesseract-Compatible Training Files

Tesseract requires .box files (character-level bounding boxes) and .lstmf files (binary LSTM training format). Since you have word-level annotations, you have two options:

  • Option A: Character-Level Training (Recommended for Nastaleeq)
    Use Tesseract to generate initial .box files, then manually correct them to match your annotations:
    tesseract your_image.jpg your_image.box -l urd makebox
    
    Open the .box file in a text editor with Urdu support (e.g., VS Code) and adjust character coordinates. Each line follows the format: [character] [x1] [y1] [x2] [y2] [page].
  • Option B: Word/Line-Level Training (Simpler)
    Wrap your word annotations in a .txt file with the same name as your image, then generate .lstmf files using the tesstrain.sh script:
    ./tesstrain.sh --fonts_dir ./fonts --fontlist "Pak Nastaleeq" "Alvi Nastaleeq" "Jameel Noori Nastaleeq" "Nafees Nastaleeq" --lang urd --linedata_only --noextract_font_properties --train_listfile train.txt --eval_listfile eval.txt
    
    (Note: train.txt and eval.txt should list paths to your image files.)

Step 3: Initialize the LSTM Model

You can either fine-tune a pre-trained model or start from scratch:

  • Fine-Tune Pre-Trained Urdu Model: Extract the LSTM component from the pre-trained urd.traineddata:
    combine_tessdata -e urd.traineddata urd.lstm
    
  • Start from Scratch: Use Tesseract's empty LSTM template:
    cp /usr/share/tesseract-ocr/4.00/tessdata/eng.lstm ./empty.lstm
    

Step 4: Train the LSTM Model

Run the training command, pointing to your data and initial model:

lstmtraining --model_output ./urdu_nastaleeq --traineddata ./tessdata/urd.traineddata --train_listfile ./train.lstmf --eval_listfile ./eval.lstmf --max_iterations 10000
  • Adjust --max_iterations based on performance (stop when evaluation loss stops decreasing).
  • Check progress periodically with:
    lstmeval --model ./urdu_nastaleeq_checkpoint --traineddata ./tessdata/urd.traineddata --eval_listfile ./eval.lstmf
    

Step 5: Finalize the Trained Model

Once training is done, generate the final .traineddata file:

lstmtraining --stop_training --continue_from ./urdu_nastaleeq_checkpoint --traineddata ./tessdata/urd.traineddata --model_output ./urdu_nastaleeq.traineddata

Test it with:

tesseract your_image.jpg output -l urdu_nastaleeq
2. Creating Word-Level Text Bounding Boxes for Urdu Images

Urdu's right-to-left script and Nastaleeq's complex ligatures require a mix of automation and manual correction. Here are your best options:

Semi-Automatic Tools (Most Efficient)

  • LabelImg: This popular annotation tool works for Urdu if you enable RTL text support in your OS. Load images, draw rectangles around each word, enter the labeled text, and save annotations in PASCAL VOC or JSON format.
  • Tesseract + Manual Correction: Generate initial boxes, then refine them:
    1. Run Tesseract to get word-level bounding boxes:
      tesseract your_image.jpg output -l urd hocr
      
    2. Extract coordinates from the .hocr file (look for <span class='ocrx_word'> tags with bbox attributes).
    3. Import these coordinates into LabelImg or a spreadsheet, then adjust misaligned boxes.

Manual Annotation (For Small Datasets)

If your dataset is small, use image editing tools to record coordinates:

  • GIMP/Photoshop: Open the image, use the rectangle selection tool to draw boxes around words, note x1,y1 (top-left) and x2,y2 (bottom-right) coordinates, then pair them with your labeled text in a CSV/JSON file.
  • Custom Spreadsheet: Create columns for image_path, x1, y1, x2, y2, text and fill in each entry manually.

Batch Processing Script (For Large Datasets)

Write a Python script to automate initial box generation, then add a manual validation step:

import cv2
import pytesseract
import json

# Configure pytesseract path and Urdu settings
pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe' # Adjust path for your OS
custom_config = r'-l urd --oem 3 --psm 6'

annotations = []
image_paths = ['image1.jpg', 'image2.jpg'] # List your image paths here

for img_path in image_paths:
    img = cv2.imread(img_path)
    # Extract word-level data
    data = pytesseract.image_to_data(img, config=custom_config, output_type=pytesseract.Output.DICT)
    num_boxes = len(data['text'])
    for i in range(num_boxes):
        if int(data['conf'][i]) > 50: # Filter low-confidence detections
            x1, y1, w, h = data['left'][i], data['top'][i], data['width'][i], data['height'][i]
            x2 = x1 + w
            y2 = y1 + h
            text = data['text'][i].strip()
            if text:
                annotations.append({
                    'image_path': img_path,
                    'bbox': [x1, y1, x2, y2],
                    'text': text
                })

# Save initial annotations for review
with open('initial_annotations.json', 'w', encoding='utf-8') as f:
    json.dump(annotations, f, ensure_ascii=False, indent=2)

After generating initial annotations, use a tool like LabelStudio to review and correct misaligned boxes or incorrect text.


内容的提问来源于stack exchange,提问作者Samee Arif

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 02:42:39