如何用Python结合Tesseract OCR自动识别文档列模式并正确读取多列内容
Got it, let's break down how to solve this mixed column OCR problem with Python and Tesseract. The issue with PSM 6 is that it treats the entire image as a single text block, so it reads horizontally across columns instead of sticking to each column sequentially. Here's a step-by-step solution that automatically detects if the document is single or double-column, then OCRs each column properly:
Your current setup uses --psm 6, which tells Tesseract to assume a single uniform text block. This fails for double-column layouts because it merges lines from left and right columns into one horizontal read. We need to:
- Detect if the image is single or double-column first
- Crop the image into individual columns
- Run OCR on each column separately
We'll use OpenCV to analyze the vertical pixel density (vertical projection) of the image. Double-column documents will have a clear low-density gap between the two columns, while single-column docs won't.
Once we know the column count and split position (if double-column), we crop each column and run Tesseract with a PSM that works for single-column text (like --psm 4, which assumes a single column of text).
import cv2 import numpy as np import pytesseract # Set Tesseract path if it's not in your system PATH # pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe' def detect_column_layout(image_path): """Detect if image is single or double-column, return column count and split x-coordinate""" # Load image and preprocess img = cv2.imread(image_path) gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # Binarize (invert so text is white, background black) _, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU) # Calculate vertical projection (sum of white pixels per column) vertical_proj = np.sum(thresh, axis=0) # Smooth projection to reduce noise smooth_proj = cv2.GaussianBlur(vertical_proj.reshape(-1, 1), (1, 7), 0).flatten() # Find low-density regions (potential column gaps) avg_density = np.mean(smooth_proj) low_density_mask = smooth_proj < avg_density * 0.1 # Adjust threshold based on your docs # Check if there's a significant continuous low-density gap consecutive_gaps = np.split(np.where(low_density_mask)[0], np.where(np.diff(np.where(low_density_mask)[0]) != 1)[0] + 1) if consecutive_gaps: largest_gap = max(consecutive_gaps, key=len) # Only count as a column split if gap is wide enough (5% of image width) if len(largest_gap) > img.shape[1] * 0.05: split_x = (largest_gap[0] + largest_gap[-1]) // 2 return 2, split_x # Default to single column return 1, None def ocr_by_column(image_path): """Run OCR on image, respecting single/double-column layout""" num_cols, split_x = detect_column_layout(image_path) img = cv2.imread(image_path) ocr_results = [] # Tesseract config: use OEM 3 (default engine) and PSM 4 (single column text) custom_config = r'--oem 3 --psm 4' if num_cols == 2: # Crop left and right columns left_col = img[:, :split_x] right_col = img[:, split_x:] # OCR each column left_text = pytesseract.image_to_string(left_col, config=custom_config).strip() right_text = pytesseract.image_to_string(right_col, config=custom_config).strip() ocr_results.append(("Left Column", left_text)) ocr_results.append(("Right Column", right_text)) else: # OCR full single column full_text = pytesseract.image_to_string(img, config=custom_config).strip() ocr_results.append(("Single Column", full_text)) return ocr_results # Example usage if __name__ == "__main__": test_image = "your_document_image.jpg" results = ocr_by_column(test_image) for col_label, text in results: print(f"--- {col_label} ---") print(text) print("\n")
- Threshold Tuning: The
avg_density * 0.1value might need tweaking if your column gaps aren't perfectly blank. Try 0.05-0.2 depending on your image quality. - PSM Selection: If cropped columns have small text blocks, switch to
--psm 6instead of4for more reliable block detection. - Preprocessing: Add tilt correction (using
pytesseract.image_to_osd()to get orientation) if your images are skewed—this will improve both column detection and OCR accuracy. - Noise Reduction: Add a Gaussian blur or morphological operation (like
cv2.morphologyEx()) before binarization if your images have a lot of noise.
内容的提问来源于stack exchange,提问作者user2910787

