You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:如何将EasyOCR检测到的文本Bounding Box提取为原图的子图像

Extracting Sub-Images from EasyOCR Bounding Boxes

Got it, let's fix that sub-image extraction part! You're already halfway there with the text detection—grabbing each bounding box as a separate sub-image is just a small, straightforward addition to your existing code. Since you're working with clean scanned documents, this should work flawlessly:

Modified Code with Sub-Image Extraction

import easyocr
import cv2
import matplotlib.pyplot as plt

IMAGE_PATH = 'test1.png'
reader = easyocr.Reader(['en'])
result = reader.readtext(IMAGE_PATH)
img = cv2.imread(IMAGE_PATH)

# Counter to keep track of sub-images (avoids overwriting files)
sub_img_index = 0

for detection in result:
    # Extract bounding box coordinates (convert to integers)
    top_left = tuple(int(val) for val in detection[0][0])
    bottom_right = tuple(int(val) for val in detection[0][2])
    
    # Break down coordinates into x/y values
    x1, y1 = top_left
    x2, y2 = bottom_right
    
    # Critical part: Slice the image to get the sub-region
    # OpenCV uses [y_start:y_end, x_start:x_end] because images are stored as (height, width)
    sub_image = img[y1:y2, x1:x2]
    
    # Save the sub-image with a unique filename
    save_path = f"extracted_text_{sub_img_index}.png"
    cv2.imwrite(save_path, sub_image)
    print(f"Saved sub-image to: {save_path}")
    
    # Optional: Keep drawing the green rectangles for visualization
    font = cv2.FONT_HERSHEY_SIMPLEX
    img = cv2.rectangle(img, top_left, bottom_right, (0,255,0), 2)
    
    sub_img_index += 1

# Fix color space for matplotlib (OpenCV reads BGR, matplotlib expects RGB)
plt.figure(figsize=(30,30))
plt.imshow(cv2.cvtColor(img, cv2.COLOR_BGR2RGB))
plt.show()

Key Notes to Avoid Headaches

  • Coordinate Order: This is likely where you got stuck! OpenCV stores images in (height, width, channels) format, so when slicing, we use [y1:y2, x1:x2] instead of the x-first order from the bounding box.
  • Unique Filenames: Using a counter (sub_img_index) ensures each sub-image gets its own file, so you won't accidentally overwrite earlier extractions.
  • Color Space Fix: Matplotlib displays images in RGB, but OpenCV loads them in BGR. Adding cv2.cvtColor(img, cv2.COLOR_BGR2RGB) makes your annotated preview look correct.

This code will save each detected text region as a separate PNG file in your working directory, while still keeping the visualization of the annotated original image if you need it.

内容的提问来源于stack exchange,提问作者Asma Alghamdi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 04:17:52