求助:如何将EasyOCR检测到的文本Bounding Box提取为原图的子图像
Extracting Sub-Images from EasyOCR Bounding Boxes
Got it, let's fix that sub-image extraction part! You're already halfway there with the text detection—grabbing each bounding box as a separate sub-image is just a small, straightforward addition to your existing code. Since you're working with clean scanned documents, this should work flawlessly:
Modified Code with Sub-Image Extraction
import easyocr import cv2 import matplotlib.pyplot as plt IMAGE_PATH = 'test1.png' reader = easyocr.Reader(['en']) result = reader.readtext(IMAGE_PATH) img = cv2.imread(IMAGE_PATH) # Counter to keep track of sub-images (avoids overwriting files) sub_img_index = 0 for detection in result: # Extract bounding box coordinates (convert to integers) top_left = tuple(int(val) for val in detection[0][0]) bottom_right = tuple(int(val) for val in detection[0][2]) # Break down coordinates into x/y values x1, y1 = top_left x2, y2 = bottom_right # Critical part: Slice the image to get the sub-region # OpenCV uses [y_start:y_end, x_start:x_end] because images are stored as (height, width) sub_image = img[y1:y2, x1:x2] # Save the sub-image with a unique filename save_path = f"extracted_text_{sub_img_index}.png" cv2.imwrite(save_path, sub_image) print(f"Saved sub-image to: {save_path}") # Optional: Keep drawing the green rectangles for visualization font = cv2.FONT_HERSHEY_SIMPLEX img = cv2.rectangle(img, top_left, bottom_right, (0,255,0), 2) sub_img_index += 1 # Fix color space for matplotlib (OpenCV reads BGR, matplotlib expects RGB) plt.figure(figsize=(30,30)) plt.imshow(cv2.cvtColor(img, cv2.COLOR_BGR2RGB)) plt.show()
Key Notes to Avoid Headaches
- Coordinate Order: This is likely where you got stuck! OpenCV stores images in
(height, width, channels)format, so when slicing, we use[y1:y2, x1:x2]instead of the x-first order from the bounding box. - Unique Filenames: Using a counter (
sub_img_index) ensures each sub-image gets its own file, so you won't accidentally overwrite earlier extractions. - Color Space Fix: Matplotlib displays images in RGB, but OpenCV loads them in BGR. Adding
cv2.cvtColor(img, cv2.COLOR_BGR2RGB)makes your annotated preview look correct.
This code will save each detected text region as a separate PNG file in your working directory, while still keeping the visualization of the annotated original image if you need it.
内容的提问来源于stack exchange,提问作者Asma Alghamdi
相关产品推荐
相关产品推荐

