如何用pytesseract与OpenCV识别被遮挡及带横线的变形文本?
Let's tackle those two annoying recognition problems you're facing—missing text near horizontal lines and failed detection when there's large content above. Here are actionable tweaks and configurations to get pytesseract working reliably:
1. Fix Missing Text Near Horizontal Lines
Horizontal lines can confuse tesseract's text detection by merging with text regions. Try these steps:
a. Preprocess the Image to Remove Lines
Use OpenCV to detect and erase horizontal lines before passing to pytesseract. This cleans up the image so tesseract focuses on text:
import cv2 import numpy as np from pytesseract import Output, image_to_data def remove_horizontal_lines(img): # Convert to grayscale gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # Create a horizontal kernel to target lines kernel = np.ones((1, 15), np.uint8) # Detect horizontal lines using morphological operations detected_lines = cv2.morphologyEx(gray, cv2.MORPH_OPEN, kernel, iterations=2) # Subtract lines from original grayscale image cleaned_img = cv2.subtract(gray, detected_lines) # Apply threshold to get a high-contrast binary image _, thresh_img = cv2.threshold(cleaned_img, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU) return thresh_img # Usage example processed_img = remove_horizontal_lines(your_input_img) td = image_to_data(processed_img, output_type=Output.DICT)
b. Adjust Tesseract's Text Detection Thresholds
Add config flags to make tesseract more sensitive to text near lines:
config = r'--oem 3 --psm 6 -c textord_min_linesize=1 textord_max_linesize=50' td = image_to_data(img, output_type=Output.DICT, config=config)
--oem 3: Uses the default hybrid engine (combines LSTM and legacy models for better compatibility)--psm 6: Treats the image as a single uniform text block (ideal for structured content like your examples)textord_min_linesize=1: Lowers the minimum line height threshold so short lines near horizontal rules aren't ignoredtextord_max_linesize=50: Sets a reasonable maximum line height to avoid merging text with large line elements
2. Fix Failed Detection of Lower Text with Large Content Above
When there's a large image or text block at the top, tesseract might skip lower regions if it thinks they're not part of the main text. Fix this with better page segmentation:
a. Use a Flexible Page Segmentation Mode
Instead of the default, use --psm 11 which tells tesseract to search for sparse text across the entire image, even with large gaps or blocks:
config = r'--oem 3 --psm 11' td = image_to_data(img, output_type=Output.DICT, config=config)
If your content is mostly structured but has large gaps, --psm 3 (fully automatic page segmentation) combined with --psm 6 can also work, but psm 11 is more reliable for scattered layouts.
b. Disable Aggressive Text Region Filtering
Add flags to prevent tesseract from discarding small or distant text regions:
config = r'--oem 3 --psm 11 -c textord_noise_rejection_threshold=0.1 textord_min_xheight=2' td = image_to_data(img, output_type=Output.DICT, config=config)
textord_noise_rejection_threshold=0.1: Reduces noise rejection so smaller text blocks aren't mistaken for noisetextord_min_xheight=2: Lowers the minimum character height threshold to catch smaller text below large elements
3. Additional Debugging Tips
- Visualize Detection Boxes: Print the bounding boxes from
image_to_datato see which regions tesseract is picking up. This helps confirm if the issue is detection or recognition:for o in range(len(td['level'])): x, y, w, h = td['left'][o], td['top'][o], td['width'][o], td['height'][o] cv2.rectangle(img, (x, y), (x+w, y+h), (0, 255, 0), 2) cv2.imshow('Detection Boxes', img) cv2.waitKey(0) - Try Adaptive Thresholding: If your image has uneven lighting, use adaptive thresholding instead of Otsu's to improve contrast:
adaptive_thresh = cv2.adaptiveThreshold(gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)
内容的提问来源于stack exchange,提问作者Navpreet Devpuri

