You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于OCR文本数据为深度学习模型准备数据集?

Preparing a Deep Learning Dataset from Receipt OCR Text

Great job getting your OCR pipeline up and running with Python 2.7! Now, turning that raw OCR text into a usable deep learning dataset involves a few key steps—let’s break them down clearly:

1. Grow Your Dataset with Diverse Receipts

Right now you’ve got one receipt’s text, but deep learning models thrive on volume and variety. Here’s how to expand:

  • Collect real receipts from different stores (grocery, office supply, retail), with varying layouts, currencies, and item types. Run your existing OCR code on each to get the text output, and keep the original image paired with the OCR result (this is useful for validating OCR accuracy later, or if you want to build an image-based model down the line).
  • If you’re short on real receipts, generate synthetic ones using tools that mimic different receipt styles (add varying fonts, text positions, and noise) then run OCR on those to bulk up your dataset.

2. Label Your OCR Data (The Most Critical Step)

Raw OCR text is unstructured—you need to tag key fields that your model will learn to identify or extract. For receipts, common labels include:

  • Merchant name (office supply hut in your example)
  • Date/time (2009-08-29 10:32 AM)
  • Cashier name (Sam)
  • Line item details: QTY, item name, price, SKU (like 1, GLUE STICK CLEARANCE, 1.99, 0476432068904)
  • Total amount (if present)

How to label efficiently:

  • Manual labeling: Start with a spreadsheet (Excel/Google Sheets) where each row is a receipt, with columns for each labeled field. For line items, use a separate sheet linked to a receipt ID to keep things organized.
  • Semi-automated labeling: Use regex to auto-tag predictable patterns (e.g., dates in YYYY-MM-DD HH:MM AM/PM format, prices with decimals). Then manually fix any OCR errors or mislabels.
  • Labeling tools: For larger datasets, use tools like LabelStudio to annotate text with custom labels quickly.

3. Structure the Dataset for Your Use Case

The structure depends on what you want your model to do:

  • Field extraction (sequence-to-sequence task): Pair raw OCR text with a structured output (like a JSON object). Example:
    {
      "raw_ocr": "office supply hut\nDate/Time: 2009-08-29 10:32 AM\nNYC DCA EL#1371013 EHASD#I47ZBB\nCashier: Sam\nQTY ITEM\n1 GLUE STICK CLEARANCE 1.99 0476432068904\n1 ROTARY CARD FILE 25.95 43...",
      "labels": {
        "merchant": "office supply hut",
        "date_time": "2009-08-29 10:32 AM",
        "cashier": "Sam",
        "line_items": [
          {"qty": 1, "item": "GLUE STICK CLEARANCE", "price": 1.99, "sku": "0476432068904"},
          {"qty": 1, "item": "ROTARY CARD FILE", "price": 25.95, "sku": "43..."}
        ]
      }
    }
    
  • Receipt classification (e.g., store type): Use a CSV with columns raw_ocr and category (e.g., "office supplies", "grocery").

4. Preprocess the OCR Text

Clean up the text to make it model-friendly:

  • Remove extra whitespace, redundant line breaks, or irrelevant characters (like random symbols from OCR errors).
  • Correct OCR mistakes (e.g., EHASD#I47ZBB might be a misread—cross-check with the original image to fix it).
  • Normalize text: convert to lowercase, standardize date formats, and remove special characters that don’t add value.
  • Tokenize: Split text into words or subwords (using tools like NLTK or Hugging Face Tokenizers) if you’re using transformer-based models.

5. Split Your Dataset

Divide your data into three sets to ensure robust training and evaluation:

  • Training set (70-80%): Used to teach the model.
  • Validation set (10-15%): Used to tune hyperparameters and catch overfitting during training.
  • Test set (10-15%): Reserved for final performance testing on unseen data.

Make sure the split is stratified—for example, if you have receipts from 10 stores, each split should have a representative sample of each store’s receipts.

6. Bonus: Pair OCR Text with Original Images

If you’re building a multi-modal model (using both text and image data), keep each OCR file linked to its original receipt image. Combining image features with text features can often improve model accuracy, especially for tricky receipts with poor OCR quality.


Content of the question来源于stack exchange,提问作者deepak patidar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:49:17