如何基于已生成的box文件训练Tesseract 3.04(适配pytesseract)?
Got it, let's walk through the exact next steps to turn your PNG/box file pairs into a custom Tesseract model that works seamlessly with pytesseract. I’ll break this down into actionable, easy-to-follow steps tailored for v3.04 (since it has different training workflows than newer Tesseract versions):
1. Organize Your Training Files First
Start by setting up a dedicated training folder (e.g., custom_ocr_train) to keep everything tidy:
- Drop all your PNG images and matching
.boxfiles here—make sure each image and box file share the exact same filename (likeinvoice01.pngandinvoice01.box). - Create a
font_propertiesfile in this folder. This tells Tesseract about the font characteristics of your training data. The format is simple:
Replacemy_custom_font 0 0 0 0 0my_custom_fontwith a name for your font (you’ll use this name later). The five 0s representitalic,bold,fixed,serif,fraktur—set them to 1 only if your font has those attributes, otherwise leave as 0.
2. Generate Training Text (.tr) Files
For each PNG/box pair, run this command to generate a .tr file (Tesseract’s core training input format):
tesseract sample_image.png sample_image nobatch box.train
If you have dozens of files, automate this with a shell script:
for img in *.png; do base=$(basename "$img" .png) tesseract "$img" "$base" nobatch box.train done
3. Create the Unicharset
The unicharset is a list of all unique characters in your training data. Generate it with:
unicharset_extractor *.box
This will create a file named unicharset in your training folder.
4. Generate Intermediate Training Files
Next, we need to build four critical files that Tesseract uses to learn character shapes and patterns. Run these commands in order, replacing my_custom_font with the name you used in font_properties:
# Cluster character shapes shapeclustering -F font_properties -U unicharset *.tr # Build font-specific unicharset and character metrics mftraining -F font_properties -U unicharset -O my_custom_font.unicharset *.tr # Generate count-based training data cntraining *.tr
After running these, you’ll have four files: inttemp, pffmtable, shapetable, and normproto.
5. Rename Intermediate Files
Rename each of these four files to include your font name—this is required for the final step:
mv inttemp my_custom_font.inttemp mv pffmtable my_custom_font.pffmtable mv shapetable my_custom_font.shapetable mv normproto my_custom_font.normproto
6. Build the Final Traineddata File
Combine all the renamed files into a single .traineddata file—this is the model pytesseract will use:
combine_tessdata my_custom_font.
Don’t forget the trailing dot! This tells combine_tessdata to look for all files prefixed with my_custom_font.. You’ll see a confirmation message if the file is created successfully (e.g., my_custom_font.traineddata).
7. Use Your Custom Model in Pytesseract
To use your new model, just specify the lang parameter with your font name when calling pytesseract:
import pytesseract from PIL import Image test_img = Image.open("your_test_image.png") extracted_text = pytesseract.image_to_string(test_img, lang="my_custom_font") print(extracted_text)
If you want to use your custom model alongside Tesseract’s default English model, separate the language codes with a plus sign: lang="eng+my_custom_font".
Quick Pro Tips
- Ensure all Tesseract training tools (
unicharset_extractor,shapeclustering, etc.) are in your system’s PATH—these usually come with a separate "training" package (e.g.,tesseract-ocr-trainingon Debian/Ubuntu). - For better accuracy, use 100+ training samples covering all characters you need to recognize.
- Double-check your box files for alignment errors—even one mispositioned character can hurt training results. Tools like jTessBoxEditor can help visualize and edit box files easily.
内容的提问来源于stack exchange,提问作者user98651

