You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于已生成的box文件训练Tesseract 3.04(适配pytesseract)?

Training Tesseract v3.04 for Pytesseract: Next Steps After PNG + Box Files

Got it, let's walk through the exact next steps to turn your PNG/box file pairs into a custom Tesseract model that works seamlessly with pytesseract. I’ll break this down into actionable, easy-to-follow steps tailored for v3.04 (since it has different training workflows than newer Tesseract versions):

1. Organize Your Training Files First

Start by setting up a dedicated training folder (e.g., custom_ocr_train) to keep everything tidy:

  • Drop all your PNG images and matching .box files here—make sure each image and box file share the exact same filename (like invoice01.png and invoice01.box).
  • Create a font_properties file in this folder. This tells Tesseract about the font characteristics of your training data. The format is simple:
    my_custom_font 0 0 0 0 0
    
    Replace my_custom_font with a name for your font (you’ll use this name later). The five 0s represent italic, bold, fixed, serif, fraktur—set them to 1 only if your font has those attributes, otherwise leave as 0.

2. Generate Training Text (.tr) Files

For each PNG/box pair, run this command to generate a .tr file (Tesseract’s core training input format):

tesseract sample_image.png sample_image nobatch box.train

If you have dozens of files, automate this with a shell script:

for img in *.png; do
    base=$(basename "$img" .png)
    tesseract "$img" "$base" nobatch box.train
done

3. Create the Unicharset

The unicharset is a list of all unique characters in your training data. Generate it with:

unicharset_extractor *.box

This will create a file named unicharset in your training folder.

4. Generate Intermediate Training Files

Next, we need to build four critical files that Tesseract uses to learn character shapes and patterns. Run these commands in order, replacing my_custom_font with the name you used in font_properties:

# Cluster character shapes
shapeclustering -F font_properties -U unicharset *.tr

# Build font-specific unicharset and character metrics
mftraining -F font_properties -U unicharset -O my_custom_font.unicharset *.tr

# Generate count-based training data
cntraining *.tr

After running these, you’ll have four files: inttemp, pffmtable, shapetable, and normproto.

5. Rename Intermediate Files

Rename each of these four files to include your font name—this is required for the final step:

mv inttemp my_custom_font.inttemp
mv pffmtable my_custom_font.pffmtable
mv shapetable my_custom_font.shapetable
mv normproto my_custom_font.normproto

6. Build the Final Traineddata File

Combine all the renamed files into a single .traineddata file—this is the model pytesseract will use:

combine_tessdata my_custom_font.

Don’t forget the trailing dot! This tells combine_tessdata to look for all files prefixed with my_custom_font.. You’ll see a confirmation message if the file is created successfully (e.g., my_custom_font.traineddata).

7. Use Your Custom Model in Pytesseract

To use your new model, just specify the lang parameter with your font name when calling pytesseract:

import pytesseract
from PIL import Image

test_img = Image.open("your_test_image.png")
extracted_text = pytesseract.image_to_string(test_img, lang="my_custom_font")
print(extracted_text)

If you want to use your custom model alongside Tesseract’s default English model, separate the language codes with a plus sign: lang="eng+my_custom_font".

Quick Pro Tips

  • Ensure all Tesseract training tools (unicharset_extractor, shapeclustering, etc.) are in your system’s PATH—these usually come with a separate "training" package (e.g., tesseract-ocr-training on Debian/Ubuntu).
  • For better accuracy, use 100+ training samples covering all characters you need to recognize.
  • Double-check your box files for alignment errors—even one mispositioned character can hurt training results. Tools like jTessBoxEditor can help visualize and edit box files easily.

内容的提问来源于stack exchange,提问作者user98651

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:35:56