You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python结合Tesseract OCR自动识别文档列模式并正确读取多列内容

Got it, let's break down how to solve this mixed column OCR problem with Python and Tesseract. The issue with PSM 6 is that it treats the entire image as a single text block, so it reads horizontally across columns instead of sticking to each column sequentially. Here's a step-by-step solution that automatically detects if the document is single or double-column, then OCRs each column properly:


Step 1: Core Problem Recap

Your current setup uses --psm 6, which tells Tesseract to assume a single uniform text block. This fails for double-column layouts because it merges lines from left and right columns into one horizontal read. We need to:

  1. Detect if the image is single or double-column first
  2. Crop the image into individual columns
  3. Run OCR on each column separately
Step 2: Automatic Column Detection with OpenCV

We'll use OpenCV to analyze the vertical pixel density (vertical projection) of the image. Double-column documents will have a clear low-density gap between the two columns, while single-column docs won't.

Step 3: Crop Columns & Run Targeted OCR

Once we know the column count and split position (if double-column), we crop each column and run Tesseract with a PSM that works for single-column text (like --psm 4, which assumes a single column of text).

Step 4: Full Code Example
import cv2
import numpy as np
import pytesseract

# Set Tesseract path if it's not in your system PATH
# pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'

def detect_column_layout(image_path):
    """Detect if image is single or double-column, return column count and split x-coordinate"""
    # Load image and preprocess
    img = cv2.imread(image_path)
    gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
    
    # Binarize (invert so text is white, background black)
    _, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)
    
    # Calculate vertical projection (sum of white pixels per column)
    vertical_proj = np.sum(thresh, axis=0)
    
    # Smooth projection to reduce noise
    smooth_proj = cv2.GaussianBlur(vertical_proj.reshape(-1, 1), (1, 7), 0).flatten()
    
    # Find low-density regions (potential column gaps)
    avg_density = np.mean(smooth_proj)
    low_density_mask = smooth_proj < avg_density * 0.1  # Adjust threshold based on your docs
    
    # Check if there's a significant continuous low-density gap
    consecutive_gaps = np.split(np.where(low_density_mask)[0], np.where(np.diff(np.where(low_density_mask)[0]) != 1)[0] + 1)
    if consecutive_gaps:
        largest_gap = max(consecutive_gaps, key=len)
        # Only count as a column split if gap is wide enough (5% of image width)
        if len(largest_gap) > img.shape[1] * 0.05:
            split_x = (largest_gap[0] + largest_gap[-1]) // 2
            return 2, split_x
    
    # Default to single column
    return 1, None

def ocr_by_column(image_path):
    """Run OCR on image, respecting single/double-column layout"""
    num_cols, split_x = detect_column_layout(image_path)
    img = cv2.imread(image_path)
    ocr_results = []
    
    # Tesseract config: use OEM 3 (default engine) and PSM 4 (single column text)
    custom_config = r'--oem 3 --psm 4'
    
    if num_cols == 2:
        # Crop left and right columns
        left_col = img[:, :split_x]
        right_col = img[:, split_x:]
        
        # OCR each column
        left_text = pytesseract.image_to_string(left_col, config=custom_config).strip()
        right_text = pytesseract.image_to_string(right_col, config=custom_config).strip()
        
        ocr_results.append(("Left Column", left_text))
        ocr_results.append(("Right Column", right_text))
    else:
        # OCR full single column
        full_text = pytesseract.image_to_string(img, config=custom_config).strip()
        ocr_results.append(("Single Column", full_text))
    
    return ocr_results

# Example usage
if __name__ == "__main__":
    test_image = "your_document_image.jpg"
    results = ocr_by_column(test_image)
    for col_label, text in results:
        print(f"--- {col_label} ---")
        print(text)
        print("\n")
Key Adjustments for Your Docs
  • Threshold Tuning: The avg_density * 0.1 value might need tweaking if your column gaps aren't perfectly blank. Try 0.05-0.2 depending on your image quality.
  • PSM Selection: If cropped columns have small text blocks, switch to --psm 6 instead of 4 for more reliable block detection.
  • Preprocessing: Add tilt correction (using pytesseract.image_to_osd() to get orientation) if your images are skewed—this will improve both column detection and OCR accuracy.
  • Noise Reduction: Add a Gaussian blur or morphological operation (like cv2.morphologyEx()) before binarization if your images have a lot of noise.

内容的提问来源于stack exchange,提问作者user2910787

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:55:01