图像处理:提取书籍目标书页并透视变换,Canny失效求解决办法
Got it, let's work through this problem together. You need to batch process book images, retaining only the page with the title Unit 3: The secret of Active Learning and discarding the other page—and it sounds like Canny edge detection didn't cut it for splitting pages. Let's switch gears to a more reliable approach that combines page segmentation with OCR text detection, since we're targeting a specific piece of text.
Why Canny Edge Detection Might Have Failed
First, let's clarify why Canny might not have worked:
- Book scans often have uneven lighting, shadow along the binding, or faint page edges that throw off edge detection.
- If the two pages are tightly aligned or have overlapping content, Canny can't reliably split them cleanly.
- Edge detection only looks for lines, not the actual content we care about—so even if it splits pages, we still need to figure out which one has our target title.
Practical Solution: OpenCV + Tesseract OCR
We'll use OpenCV to handle image preprocessing and page splitting, then Tesseract OCR to read text from each split page and check for our target title. Here's a reusable, batch-ready code setup:
Step 1: Install Dependencies
First, install the required libraries and Tesseract engine:
pip install opencv-python pytesseract
(Note: You'll also need to install Tesseract itself on your system—follow the official docs for your OS if you haven't already.)
Step 2: Batch Processing Code
import cv2 import pytesseract import os # Configure Tesseract path if it's not in your system PATH (adjust for your OS) # pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe' # Windows example # pytesseract.pytesseract.tesseract_cmd = '/usr/bin/tesseract' # Linux/macOS example TARGET_TITLE = "Unit 3: The secret of Active Learning" def process_book_image(input_path, output_dir): # Load image and convert to grayscale img = cv2.imread(input_path) gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # Preprocess for better OCR: apply thresholding to reduce noise _, thresh = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU) # Split the image into left and right pages (adjust split ratio if needed) height, width = thresh.shape mid_point = width // 2 left_page = thresh[:, :mid_point] right_page = thresh[:, mid_point:] # Function to check if a page contains the target title def has_target_title(page_img): # Use Tesseract to extract text (adjust config for better accuracy) custom_config = r'--oem 3 --psm 6' # OEM 3 = default engine, PSM 6 = assume a single uniform block of text text = pytesseract.image_to_string(page_img, config=custom_config) return TARGET_TITLE in text # Determine which page to keep if has_target_title(left_page): final_page = img[:, :mid_point] elif has_target_title(right_page): final_page = img[:, mid_point:] else: print(f"Warning: Target title not found in {input_path}") return # Save the processed image filename = os.path.basename(input_path) output_path = os.path.join(output_dir, filename) cv2.imwrite(output_path, final_page) print(f"Processed and saved: {output_path}") def batch_process_images(input_dir, output_dir): # Create output directory if it doesn't exist os.makedirs(output_dir, exist_ok=True) # Iterate over all image files in input directory for filename in os.listdir(input_dir): if filename.lower().endswith(('.png', '.jpg', '.jpeg', '.bmp')): input_path = os.path.join(input_dir, filename) process_book_image(input_path, output_dir) # Example usage if __name__ == "__main__": INPUT_DIRECTORY = "path/to/your/book_images" OUTPUT_DIRECTORY = "path/to/save/processed_images" batch_process_images(INPUT_DIRECTORY, OUTPUT_DIRECTORY)
Key Customization Tips
- Adjust Page Split: If your images have pages stacked vertically instead of horizontally, swap the split logic to divide by height instead of width.
- OCR Accuracy Tweaks: Modify the
custom_configparameter for Tesseract—trypsm 3(default page segmentation) ifpsm 6doesn't work, or add language flags if your text isn't in English. - Preprocessing Adjustments: If OCR is struggling, add a blur step before thresholding:
blurred = cv2.GaussianBlur(gray, (5,5), 0)then useblurredfor thresholding. - Handle Rotated Pages: If some images are rotated, add a step to detect and correct orientation using Tesseract's orientation detection.
Troubleshooting
- If Tesseract can't find the title, check the preprocessing steps—make sure the text is clear and contrast is high.
- For images with non-standard page layouts (e.g., more than two pages), you'll need to adjust the segmentation logic (like using contour detection to find individual page boundaries).
内容的提问来源于stack exchange,提问作者Ninad Naik

