乌尔都语Tesseract 4.00训练流程及图像单词标注方法咨询
Great to hear you've got Tesseract 4.0 set up and a solid Urdu dataset with those popular Nastaleeq fonts—those are tricky but super useful for training a robust model! Let's break down your two questions step by step:
Tesseract 4.0 uses an LSTM-based model, so we'll focus on fine-tuning (or building from scratch) a model optimized for your Nastaleeq fonts. Here's the full process tailored to your setup:
Step 1: Preprocess Your Dataset
- Image Normalization: Ensure all images have consistent resolution (e.g., 300 DPI), clean up noise (using tools like OpenCV's
bilateralFilter), and confirm text is correctly oriented (Urdu is right-to-left, so no upside-down or rotated images). - Clean Annotations: Double-check your labeled text to remove extra spaces, typos, or non-Urdu characters. For Nastaleeq fonts, pay close attention to ligatures (joined characters) — your annotations should match the exact rendered text (e.g.,
بعد نجی ٹی ویinstead of incorrectly splitting connected characters). - Split Data: Divide your dataset into training (80-90%) and evaluation (10-20%) sets to monitor model performance during training.
Step 2: Generate Tesseract-Compatible Training Files
Tesseract requires .box files (character-level bounding boxes) and .lstmf files (binary LSTM training format). Since you have word-level annotations, you have two options:
- Option A: Character-Level Training (Recommended for Nastaleeq)
Use Tesseract to generate initial.boxfiles, then manually correct them to match your annotations:
Open thetesseract your_image.jpg your_image.box -l urd makebox.boxfile in a text editor with Urdu support (e.g., VS Code) and adjust character coordinates. Each line follows the format:[character] [x1] [y1] [x2] [y2] [page]. - Option B: Word/Line-Level Training (Simpler)
Wrap your word annotations in a.txtfile with the same name as your image, then generate.lstmffiles using thetesstrain.shscript:
(Note:./tesstrain.sh --fonts_dir ./fonts --fontlist "Pak Nastaleeq" "Alvi Nastaleeq" "Jameel Noori Nastaleeq" "Nafees Nastaleeq" --lang urd --linedata_only --noextract_font_properties --train_listfile train.txt --eval_listfile eval.txttrain.txtandeval.txtshould list paths to your image files.)
Step 3: Initialize the LSTM Model
You can either fine-tune a pre-trained model or start from scratch:
- Fine-Tune Pre-Trained Urdu Model: Extract the LSTM component from the pre-trained
urd.traineddata:combine_tessdata -e urd.traineddata urd.lstm - Start from Scratch: Use Tesseract's empty LSTM template:
cp /usr/share/tesseract-ocr/4.00/tessdata/eng.lstm ./empty.lstm
Step 4: Train the LSTM Model
Run the training command, pointing to your data and initial model:
lstmtraining --model_output ./urdu_nastaleeq --traineddata ./tessdata/urd.traineddata --train_listfile ./train.lstmf --eval_listfile ./eval.lstmf --max_iterations 10000
- Adjust
--max_iterationsbased on performance (stop when evaluation loss stops decreasing). - Check progress periodically with:
lstmeval --model ./urdu_nastaleeq_checkpoint --traineddata ./tessdata/urd.traineddata --eval_listfile ./eval.lstmf
Step 5: Finalize the Trained Model
Once training is done, generate the final .traineddata file:
lstmtraining --stop_training --continue_from ./urdu_nastaleeq_checkpoint --traineddata ./tessdata/urd.traineddata --model_output ./urdu_nastaleeq.traineddata
Test it with:
tesseract your_image.jpg output -l urdu_nastaleeq
Urdu's right-to-left script and Nastaleeq's complex ligatures require a mix of automation and manual correction. Here are your best options:
Semi-Automatic Tools (Most Efficient)
- LabelImg: This popular annotation tool works for Urdu if you enable RTL text support in your OS. Load images, draw rectangles around each word, enter the labeled text, and save annotations in PASCAL VOC or JSON format.
- Tesseract + Manual Correction: Generate initial boxes, then refine them:
- Run Tesseract to get word-level bounding boxes:
tesseract your_image.jpg output -l urd hocr - Extract coordinates from the
.hocrfile (look for<span class='ocrx_word'>tags withbboxattributes). - Import these coordinates into LabelImg or a spreadsheet, then adjust misaligned boxes.
- Run Tesseract to get word-level bounding boxes:
Manual Annotation (For Small Datasets)
If your dataset is small, use image editing tools to record coordinates:
- GIMP/Photoshop: Open the image, use the rectangle selection tool to draw boxes around words, note x1,y1 (top-left) and x2,y2 (bottom-right) coordinates, then pair them with your labeled text in a CSV/JSON file.
- Custom Spreadsheet: Create columns for
image_path,x1,y1,x2,y2,textand fill in each entry manually.
Batch Processing Script (For Large Datasets)
Write a Python script to automate initial box generation, then add a manual validation step:
import cv2 import pytesseract import json # Configure pytesseract path and Urdu settings pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe' # Adjust path for your OS custom_config = r'-l urd --oem 3 --psm 6' annotations = [] image_paths = ['image1.jpg', 'image2.jpg'] # List your image paths here for img_path in image_paths: img = cv2.imread(img_path) # Extract word-level data data = pytesseract.image_to_data(img, config=custom_config, output_type=pytesseract.Output.DICT) num_boxes = len(data['text']) for i in range(num_boxes): if int(data['conf'][i]) > 50: # Filter low-confidence detections x1, y1, w, h = data['left'][i], data['top'][i], data['width'][i], data['height'][i] x2 = x1 + w y2 = y1 + h text = data['text'][i].strip() if text: annotations.append({ 'image_path': img_path, 'bbox': [x1, y1, x2, y2], 'text': text }) # Save initial annotations for review with open('initial_annotations.json', 'w', encoding='utf-8') as f: json.dump(annotations, f, ensure_ascii=False, indent=2)
After generating initial annotations, use a tool like LabelStudio to review and correct misaligned boxes or incorrect text.
内容的提问来源于stack exchange,提问作者Samee Arif

