如何使用Mindee Doctr提取图像特定区域的文本?
提取Doctr OCR结果中特定区域文本的可行方案
下面是几种实用的实现方案,可根据你的场景选择:
方案1:裁剪目标区域后再执行OCR
如果已经明确目标区域的像素坐标,直接裁剪图像后再用Doctr识别,是最直观且准确率较高的方式:
from PIL import Image from doctr.models import ocr_predictor from doctr.io import DocumentFile # 打开原始图像并裁剪目标区域 img = Image.open("docs/temp.jpg") # 裁剪区域格式:(左边界像素, 上边界像素, 右边界像素, 下边界像素),根据实际需求调整 crop_area = (100, 200, 500, 600) cropped_img = img.crop(crop_area) cropped_img.save("docs/cropped_temp.jpg") # 对裁剪后的图像执行OCR model = ocr_predictor(reco_arch='crnn_vgg16_bn', pretrained=True, export_as_straight_boxes=True) image = DocumentFile.from_images("docs/cropped_temp.jpg") result = model(image) # 提取并拼接文本 extracted_text = "\n".join([word.value for page in result.pages for block in page.blocks for line in block.lines for word in line.words]) print(extracted_text)
方案2:解析OCR结果的坐标筛选文本
Doctr的OCR结果会返回每个文字/文本块的归一化坐标(范围0-1,相对于图像宽高),可以基于这些坐标过滤出目标区域内的文本,无需重新处理图像:
from doctr.models import ocr_predictor from doctr.io import DocumentFile model = ocr_predictor(reco_arch='crnn_vgg16_bn', pretrained=True, export_as_straight_boxes=True) image = DocumentFile.from_images("docs/temp.jpg") result = model(image) result_json = result.export() # 定义目标区域的归一化坐标:(x_min, y_min, x_max, y_max) # 若你只有像素坐标,需除以图像宽高转换:比如x_min = 像素左边界 / 图像宽度 target_area = (0.1, 0.2, 0.5, 0.6) x_min, y_min, x_max, y_max = target_area extracted_text = [] # 遍历所有文本元素,筛选坐标在目标区域内的内容 for page in result_json["pages"]: for block in page["blocks"]: for line in block["lines"]: for word in line["words"]: # 获取文字的 bounding box 并计算中心坐标 word_box = word["geometry"] word_center_x = (word_box[0][0] + word_box[1][0]) / 2 word_center_y = (word_box[0][1] + word_box[1][1]) / 2 # 判断中心坐标是否在目标区域内 if x_min <= word_center_x <= x_max and y_min <= word_center_y <= y_max: extracted_text.append(word["value"]) # 拼接最终文本 final_text = " ".join(extracted_text) print(final_text)
方案3:先检测文本块再针对性识别
如果无法确定精确坐标,但能通过文本块的位置/顺序定位目标区域,可以先单独调用检测模型定位文本块,再裁剪识别:
from doctr.models import detection_predictor, recognition_predictor from doctr.io import DocumentFile from PIL import Image import numpy as np # 加载检测和识别模型 det_model = detection_predictor(pretrained=True) rec_model = recognition_predictor(reco_arch='crnn_vgg16_bn', pretrained=True) # 检测所有文本块 image = DocumentFile.from_images("docs/temp.jpg") det_result = det_model(image) det_json = det_result.export() # 定位目标文本块(示例为第2个块,根据实际情况调整索引) target_block = det_json["pages"][0]["blocks"][1] block_box = target_block["geometry"] # 将归一化坐标转换为像素坐标 img_pil = Image.open("docs/temp.jpg") img_width, img_height = img_pil.size x1 = int(block_box[0][0] * img_width) y1 = int(block_box[0][1] * img_height) x2 = int(block_box[1][0] * img_width) y2 = int(block_box[1][1] * img_height) # 裁剪目标块并识别 cropped_img = img_pil.crop((x1, y1, x2, y2)) cropped_doc = DocumentFile.from_images(np.array(cropped_img)) rec_result = rec_model(cropped_doc) rec_json = rec_result.export() # 提取文本 block_text = "\n".join([word["value"] for line in rec_json["pages"][0]["blocks"][0]["lines"] for word in line["words"]]) print(block_text)
内容的提问来源于stack exchange,提问作者RAVINDER SINGH
相关产品推荐
相关产品推荐

