如何基于doctr识别的坐标绘制点间连线并实现文档模板检测?
解决方案:从识别坐标生成图形并批量检测文档
一、基于识别坐标生成模板图形
你拿到的是相对坐标(每组为目标区域的左上角(x1,y1)和右下角(x2,y2)),先处理坐标再绘图:
1. 坐标预处理
先提取每个目标区域的中心坐标(作为连线的关键点):
# 输入的三组doctr识别坐标(相对比例) bboxes = [ ((0.09370404411764705, 0.0439453125), (0.33140099789915967, 0.09765625)), ((0.5925912552521009, 0.1796875), (0.6575433298319328, 0.1953125)), ((0.5925912552521009, 0.2041015625), (0.6575433298319328, 0.21875)) ] # 计算每个区域的中心坐标 centers = [] for (x1, y1), (x2, y2) in bboxes: cx = (x1 + x2) / 2 cy = (y1 + y2) / 2 centers.append((cx, cy))
2. 绘制连线图形
用matplotlib生成模板图形(也可生成像素图用于后续匹配):
import matplotlib.pyplot as plt # 模拟A4文档比例的画布 fig, ax = plt.subplots(figsize=(7, 10)) # 标记每个中心 for cx, cy in centers: ax.scatter(cx, cy, color='red', s=50) # 按顺序连接中心 ax.plot([p[0] for p in centers], [p[1] for p in centers], color='blue', linewidth=2) # 隐藏坐标轴,只保留图形 ax.axis('off') plt.savefig('invoice_template.png', bbox_inches='tight', pad_inches=0, dpi=300) plt.close()
生成的invoice_template.png就是包含目标点连线结构的模板。
二、批量检测文档中的目标图形
推荐两种可靠的检测方式:
方式1:坐标相对位置比对(不受缩放影响,优先选择)
无需模板图片,直接用doctr识别批量文档的目标单词坐标,比对相对位置是否匹配模板:
from doctr.io import DocumentFile from doctr.models import ocr_predictor import os # 加载预训练OCR模型 model = ocr_predictor(pretrained=True) def check_template_match(doc_path, template_centers): # 识别文档中的目标单词坐标 doc = DocumentFile.from_pdf(doc_path) result = model(doc) detected_centers = [] for page in result.pages: for block in page.blocks: for line in block.lines: for word in line.words: if word.value == "INVOICE": # 匹配目标单词 x1, y1 = word.geometry[0] x2, y2 = word.geometry[1] cx = (x1 + x2)/2 cy = (y1 + y2)/2 detected_centers.append((cx, cy)) # 检查点数量是否匹配 if len(detected_centers) != len(template_centers): return False # 计算模板的相对位置比例 base_x, base_y = template_centers[0] template_ratios = [((cx - base_x)/base_x, (cy - base_y)/base_y) for cx, cy in template_centers[1:]] # 计算检测点的相对位置比例 base_dx, base_dy = detected_centers[0] detected_ratios = [((cx - base_dx)/base_dx, (cy - base_dy)/base_dy) for cx, cy in detected_centers[1:]] # 允许5%的误差范围 for t_ratio, d_ratio in zip(template_ratios, detected_ratios): if abs(t_ratio[0] - d_ratio[0]) > 0.05 or abs(t_ratio[1] - d_ratio[1]) > 0.05: return False return True # 批量处理示例 target_dir = "/path/to/your/pdf/files" for filename in os.listdir(target_dir): if filename.endswith(".pdf"): file_path = os.path.join(target_dir, filename) if check_template_match(file_path, centers): print(f"符合条件的文档:{filename}")
方式2:模板图片匹配(适合固定尺寸文档)
如果文档尺寸统一,用OpenCV做模板匹配:
import cv2 import numpy as np from doctr.io import DocumentFile import os # 加载模板 template = cv2.imread('invoice_template.png', 0) w, h = template.shape[::-1] def match_template(doc_path): # 将PDF第一页转为图片 doc = DocumentFile.from_pdf(doc_path) page_np = np.array(doc[0]) gray = cv2.cvtColor(page_np, cv2.COLOR_RGB2GRAY) # 模板匹配 res = cv2.matchTemplate(gray, template, cv2.TM_CCOEFF_NORMED) threshold = 0.8 # 匹配阈值,可根据实际调整 loc = np.where(res >= threshold) return len(loc[0]) > 0 # 批量处理示例 target_dir = "/path/to/your/pdf/files" for filename in os.listdir(target_dir): if filename.endswith(".pdf"): file_path = os.path.join(target_dir, filename) if match_template(file_path): print(f"符合条件的文档:{filename}")
关键注意事项
- 相对坐标优势:doctr返回的是页面比例坐标,不受文档缩放影响,用位置比例比对更鲁棒。
- 目标过滤:确保批量识别时准确匹配"INVOICE",避免误抓其他单词的坐标。
- 阈值调整:比例误差阈值、模板匹配阈值需根据实际文档调整,平衡准确率和召回率。
内容的提问来源于stack exchange,提问作者Paul
相关产品推荐
相关产品推荐

