You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于doctr识别的坐标绘制点间连线并实现文档模板检测?

解决方案:从识别坐标生成图形并批量检测文档

一、基于识别坐标生成模板图形

你拿到的是相对坐标(每组为目标区域的左上角(x1,y1)和右下角(x2,y2)),先处理坐标再绘图:

1. 坐标预处理

先提取每个目标区域的中心坐标(作为连线的关键点):

# 输入的三组doctr识别坐标(相对比例)
bboxes = [
    ((0.09370404411764705, 0.0439453125), (0.33140099789915967, 0.09765625)),
    ((0.5925912552521009, 0.1796875), (0.6575433298319328, 0.1953125)),
    ((0.5925912552521009, 0.2041015625), (0.6575433298319328, 0.21875))
]

# 计算每个区域的中心坐标
centers = []
for (x1, y1), (x2, y2) in bboxes:
    cx = (x1 + x2) / 2
    cy = (y1 + y2) / 2
    centers.append((cx, cy))

2. 绘制连线图形

用matplotlib生成模板图形(也可生成像素图用于后续匹配):

import matplotlib.pyplot as plt

# 模拟A4文档比例的画布
fig, ax = plt.subplots(figsize=(7, 10))

# 标记每个中心
for cx, cy in centers:
    ax.scatter(cx, cy, color='red', s=50)

# 按顺序连接中心
ax.plot([p[0] for p in centers], [p[1] for p in centers], color='blue', linewidth=2)

# 隐藏坐标轴,只保留图形
ax.axis('off')
plt.savefig('invoice_template.png', bbox_inches='tight', pad_inches=0, dpi=300)
plt.close()

生成的invoice_template.png就是包含目标点连线结构的模板。

二、批量检测文档中的目标图形

推荐两种可靠的检测方式:

方式1:坐标相对位置比对(不受缩放影响,优先选择)

无需模板图片,直接用doctr识别批量文档的目标单词坐标,比对相对位置是否匹配模板:

from doctr.io import DocumentFile
from doctr.models import ocr_predictor
import os

# 加载预训练OCR模型
model = ocr_predictor(pretrained=True)

def check_template_match(doc_path, template_centers):
    # 识别文档中的目标单词坐标
    doc = DocumentFile.from_pdf(doc_path)
    result = model(doc)
    
    detected_centers = []
    for page in result.pages:
        for block in page.blocks:
            for line in block.lines:
                for word in line.words:
                    if word.value == "INVOICE":  # 匹配目标单词
                        x1, y1 = word.geometry[0]
                        x2, y2 = word.geometry[1]
                        cx = (x1 + x2)/2
                        cy = (y1 + y2)/2
                        detected_centers.append((cx, cy))
    
    # 检查点数量是否匹配
    if len(detected_centers) != len(template_centers):
        return False
    
    # 计算模板的相对位置比例
    base_x, base_y = template_centers[0]
    template_ratios = [((cx - base_x)/base_x, (cy - base_y)/base_y) for cx, cy in template_centers[1:]]
    
    # 计算检测点的相对位置比例
    base_dx, base_dy = detected_centers[0]
    detected_ratios = [((cx - base_dx)/base_dx, (cy - base_dy)/base_dy) for cx, cy in detected_centers[1:]]
    
    # 允许5%的误差范围
    for t_ratio, d_ratio in zip(template_ratios, detected_ratios):
        if abs(t_ratio[0] - d_ratio[0]) > 0.05 or abs(t_ratio[1] - d_ratio[1]) > 0.05:
            return False
    return True

# 批量处理示例
target_dir = "/path/to/your/pdf/files"
for filename in os.listdir(target_dir):
    if filename.endswith(".pdf"):
        file_path = os.path.join(target_dir, filename)
        if check_template_match(file_path, centers):
            print(f"符合条件的文档:{filename}")

方式2:模板图片匹配(适合固定尺寸文档)

如果文档尺寸统一,用OpenCV做模板匹配:

import cv2
import numpy as np
from doctr.io import DocumentFile
import os

# 加载模板
template = cv2.imread('invoice_template.png', 0)
w, h = template.shape[::-1]

def match_template(doc_path):
    # 将PDF第一页转为图片
    doc = DocumentFile.from_pdf(doc_path)
    page_np = np.array(doc[0])
    gray = cv2.cvtColor(page_np, cv2.COLOR_RGB2GRAY)
    
    # 模板匹配
    res = cv2.matchTemplate(gray, template, cv2.TM_CCOEFF_NORMED)
    threshold = 0.8  # 匹配阈值,可根据实际调整
    loc = np.where(res >= threshold)
    
    return len(loc[0]) > 0

# 批量处理示例
target_dir = "/path/to/your/pdf/files"
for filename in os.listdir(target_dir):
    if filename.endswith(".pdf"):
        file_path = os.path.join(target_dir, filename)
        if match_template(file_path):
            print(f"符合条件的文档:{filename}")

关键注意事项

  • 相对坐标优势:doctr返回的是页面比例坐标,不受文档缩放影响,用位置比例比对更鲁棒。
  • 目标过滤:确保批量识别时准确匹配"INVOICE",避免误抓其他单词的坐标。
  • 阈值调整:比例误差阈值、模板匹配阈值需根据实际文档调整,平衡准确率和召回率。

内容的提问来源于stack exchange,提问作者Paul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 22:32:09