如何利用膨胀图像掩码提取发票收据中的表格数据?
Extracting Table Data from Invoice/Receipt Images Using Inflated Mask
Got it, let's walk through how to turn that inflated image mask you've got into structured table data—covering both bordered and borderless tables from invoices/receipts. Here's a practical, step-by-step approach:
1. Refine the Inflated Mask First
Since it's an inflated mask, you might have over-expanded regions that merge adjacent cells (especially tricky for borderless tables). Fix this with quick morphological tweaks:
- Use erosion to shrink the mask slightly and separate merged cells:
import cv2 import numpy as np kernel = np.ones((2,2), np.uint8) refined_mask = cv2.erode(inflated_mask, kernel, iterations=1) - Clear small noise blobs with an opening operation:
refined_mask = cv2.morphologyEx(refined_mask, cv2.MORPH_OPEN, kernel, iterations=1)
2. Locate and Segment Individual Cells
Next, you need to map each cell's boundary and split them from the original image:
- Extract contours from the refined mask (focus on external contours only):
contours, _ = cv2.findContours(refined_mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) - Filter out tiny noise contours and sort the rest into logical rows/columns:
- Group contours by their y-coordinate (add a small tolerance for minor vertical misalignment) to form rows
- Within each row, sort contours by x-coordinate to get left-to-right column order
- For bordered tables, contours will directly match cell boundaries; for borderless tables, the inflated mask should already highlight each cell's region, so contours will define those areas perfectly.
3. Extract Text from Each Cell
Once cells are segmented, use OCR to pull out the text with better accuracy:
- Crop the original image to each cell's bounding box:
cell_imgs = [] for cnt in sorted_contours: x, y, w, h = cv2.boundingRect(cnt) cell_img = original_img[y:y+h, x:x+w] cell_imgs.append(cell_img) - Preprocess cell images to boost OCR performance:
def preprocess_cell(img): gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) _, binary = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU) return binary - Use Tesseract with a PSM mode optimized for single-cell text:
import pytesseract # Windows users: uncomment and set your Tesseract path # pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe' extracted_text = [pytesseract.image_to_string(preprocess_cell(img), config='--psm 6').strip() for img in cell_imgs]
4. Structure the Extracted Data
Finally, turn the flat text list into a usable structured format:
- Reshape the text into a 2D list matching the table's row-column layout:
import pandas as pd # Calculate number of columns from your sorted contours (e.g., count of contours in first row) num_cols = len(sorted_contours[0]) if isinstance(sorted_contours[0], list) else 4 # Replace with your actual column count table_data = [extracted_text[i:i+num_cols] for i in range(0, len(extracted_text), num_cols)] df = pd.DataFrame(table_data) - For borderless tables, double-check row grouping—adjust the y-coordinate tolerance during contour sorting if lines are misaligned.
Quick Pro Tips
- If your original images are skewed, use the table's detection contours to apply a perspective transform first
- Add a deskew step for low-quality images to further improve OCR accuracy
- Test with both bordered and borderless samples to tweak morphological parameters (kernel size, iteration count)
内容的提问来源于stack exchange,提问作者Archit
相关产品推荐
相关产品推荐

