You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Mindee Doctr提取图像特定区域的文本?

提取Doctr OCR结果中特定区域文本的可行方案

下面是几种实用的实现方案,可根据你的场景选择:

方案1:裁剪目标区域后再执行OCR

如果已经明确目标区域的像素坐标,直接裁剪图像后再用Doctr识别,是最直观且准确率较高的方式:

from PIL import Image
from doctr.models import ocr_predictor
from doctr.io import DocumentFile

# 打开原始图像并裁剪目标区域
img = Image.open("docs/temp.jpg")
# 裁剪区域格式:(左边界像素, 上边界像素, 右边界像素, 下边界像素),根据实际需求调整
crop_area = (100, 200, 500, 600)
cropped_img = img.crop(crop_area)
cropped_img.save("docs/cropped_temp.jpg")

# 对裁剪后的图像执行OCR
model = ocr_predictor(reco_arch='crnn_vgg16_bn', pretrained=True, export_as_straight_boxes=True)
image = DocumentFile.from_images("docs/cropped_temp.jpg")
result = model(image)

# 提取并拼接文本
extracted_text = "\n".join([word.value for page in result.pages for block in page.blocks for line in block.lines for word in line.words])
print(extracted_text)

方案2:解析OCR结果的坐标筛选文本

Doctr的OCR结果会返回每个文字/文本块的归一化坐标(范围0-1,相对于图像宽高),可以基于这些坐标过滤出目标区域内的文本,无需重新处理图像:

from doctr.models import ocr_predictor
from doctr.io import DocumentFile

model = ocr_predictor(reco_arch='crnn_vgg16_bn', pretrained=True, export_as_straight_boxes=True)
image = DocumentFile.from_images("docs/temp.jpg")
result = model(image)
result_json = result.export()

# 定义目标区域的归一化坐标:(x_min, y_min, x_max, y_max)
# 若你只有像素坐标,需除以图像宽高转换:比如x_min = 像素左边界 / 图像宽度
target_area = (0.1, 0.2, 0.5, 0.6)
x_min, y_min, x_max, y_max = target_area

extracted_text = []
# 遍历所有文本元素,筛选坐标在目标区域内的内容
for page in result_json["pages"]:
    for block in page["blocks"]:
        for line in block["lines"]:
            for word in line["words"]:
                # 获取文字的 bounding box 并计算中心坐标
                word_box = word["geometry"]
                word_center_x = (word_box[0][0] + word_box[1][0]) / 2
                word_center_y = (word_box[0][1] + word_box[1][1]) / 2
                # 判断中心坐标是否在目标区域内
                if x_min <= word_center_x <= x_max and y_min <= word_center_y <= y_max:
                    extracted_text.append(word["value"])

# 拼接最终文本
final_text = " ".join(extracted_text)
print(final_text)

方案3:先检测文本块再针对性识别

如果无法确定精确坐标,但能通过文本块的位置/顺序定位目标区域,可以先单独调用检测模型定位文本块,再裁剪识别:

from doctr.models import detection_predictor, recognition_predictor
from doctr.io import DocumentFile
from PIL import Image
import numpy as np

# 加载检测和识别模型
det_model = detection_predictor(pretrained=True)
rec_model = recognition_predictor(reco_arch='crnn_vgg16_bn', pretrained=True)

# 检测所有文本块
image = DocumentFile.from_images("docs/temp.jpg")
det_result = det_model(image)
det_json = det_result.export()

# 定位目标文本块(示例为第2个块,根据实际情况调整索引)
target_block = det_json["pages"][0]["blocks"][1]
block_box = target_block["geometry"]

# 将归一化坐标转换为像素坐标
img_pil = Image.open("docs/temp.jpg")
img_width, img_height = img_pil.size
x1 = int(block_box[0][0] * img_width)
y1 = int(block_box[0][1] * img_height)
x2 = int(block_box[1][0] * img_width)
y2 = int(block_box[1][1] * img_height)

# 裁剪目标块并识别
cropped_img = img_pil.crop((x1, y1, x2, y2))
cropped_doc = DocumentFile.from_images(np.array(cropped_img))
rec_result = rec_model(cropped_doc)
rec_json = rec_result.export()

# 提取文本
block_text = "\n".join([word["value"] for line in rec_json["pages"][0]["blocks"][0]["lines"] for word in line["words"]])
print(block_text)

内容的提问来源于stack exchange,提问作者RAVINDER SINGH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 21:57:44