You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用膨胀图像掩码提取发票收据中的表格数据?

Extracting Table Data from Invoice/Receipt Images Using Inflated Mask

Got it, let's walk through how to turn that inflated image mask you've got into structured table data—covering both bordered and borderless tables from invoices/receipts. Here's a practical, step-by-step approach:

1. Refine the Inflated Mask First

Since it's an inflated mask, you might have over-expanded regions that merge adjacent cells (especially tricky for borderless tables). Fix this with quick morphological tweaks:

  • Use erosion to shrink the mask slightly and separate merged cells:
    import cv2
    import numpy as np
    
    kernel = np.ones((2,2), np.uint8)
    refined_mask = cv2.erode(inflated_mask, kernel, iterations=1)
    
  • Clear small noise blobs with an opening operation:
    refined_mask = cv2.morphologyEx(refined_mask, cv2.MORPH_OPEN, kernel, iterations=1)
    

2. Locate and Segment Individual Cells

Next, you need to map each cell's boundary and split them from the original image:

  • Extract contours from the refined mask (focus on external contours only):
    contours, _ = cv2.findContours(refined_mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
    
  • Filter out tiny noise contours and sort the rest into logical rows/columns:
    • Group contours by their y-coordinate (add a small tolerance for minor vertical misalignment) to form rows
    • Within each row, sort contours by x-coordinate to get left-to-right column order
    • For bordered tables, contours will directly match cell boundaries; for borderless tables, the inflated mask should already highlight each cell's region, so contours will define those areas perfectly.

3. Extract Text from Each Cell

Once cells are segmented, use OCR to pull out the text with better accuracy:

  • Crop the original image to each cell's bounding box:
    cell_imgs = []
    for cnt in sorted_contours:
        x, y, w, h = cv2.boundingRect(cnt)
        cell_img = original_img[y:y+h, x:x+w]
        cell_imgs.append(cell_img)
    
  • Preprocess cell images to boost OCR performance:
    def preprocess_cell(img):
        gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
        _, binary = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)
        return binary
    
  • Use Tesseract with a PSM mode optimized for single-cell text:
    import pytesseract
    
    # Windows users: uncomment and set your Tesseract path
    # pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'
    extracted_text = [pytesseract.image_to_string(preprocess_cell(img), config='--psm 6').strip() for img in cell_imgs]
    

4. Structure the Extracted Data

Finally, turn the flat text list into a usable structured format:

  • Reshape the text into a 2D list matching the table's row-column layout:
    import pandas as pd
    
    # Calculate number of columns from your sorted contours (e.g., count of contours in first row)
    num_cols = len(sorted_contours[0]) if isinstance(sorted_contours[0], list) else 4 # Replace with your actual column count
    table_data = [extracted_text[i:i+num_cols] for i in range(0, len(extracted_text), num_cols)]
    df = pd.DataFrame(table_data)
    
  • For borderless tables, double-check row grouping—adjust the y-coordinate tolerance during contour sorting if lines are misaligned.

Quick Pro Tips

  • If your original images are skewed, use the table's detection contours to apply a perspective transform first
  • Add a deskew step for low-quality images to further improve OCR accuracy
  • Test with both bordered and borderless samples to tweak morphological parameters (kernel size, iteration count)

内容的提问来源于stack exchange,提问作者Archit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:23:25