如何使用Python提取PDF中各元素(文本、图片、表格)的坐标
PDF文档元素坐标提取方案
这个需求完全可以实现,你当前使用传统OpenCV轮廓检测+Tesseract的方案效果不佳,核心问题是规则化的传统视觉方法无法适配复杂排版的文档,也不能对文本、图片、表格三类元素做分类识别。
现有方案存在的问题
- 纯OpenCV轮廓检测仅能识别边界特征,无法区分轮廓对应的元素类型,很容易把文本行、表格边框、图片边缘混判,也无法自动合并属于同一文本块的分散轮廓
- 单独使用Tesseract时如果没有做前置区域分割,会出现文本块拆分错误、坐标偏移的问题,也无法输出表格、图片的坐标信息
可落地的实现方案
轻量原生PDF适配方案(无需训练,仅适用于可复制的非扫描PDF)
不需要将PDF转成图片用CV处理,直接调用PDF解析工具即可精准获取所有元素坐标:
- 文本块:使用PyMuPDF库的
page.get_text("blocks")接口,返回结果直接包含每个文本块的左上角、右下角坐标以及对应文本内容 - 图片:使用PyMuPDF库的
page.get_images()接口,可直接获取每个插入图片的坐标、大小与原始数据 - 表格:配合camelot或tabula-py工具提取表格内容的同时,即可获取对应表格的页面坐标,识别准确率远高于OpenCV边框检测
深度学习通用方案(适配扫描版PDF、复杂排版文档)
你示例中用到的目标检测方案是当前通用场景下准确率最高的方案,不需要从零开发训练,直接使用成熟预训练模型即可:
- 可直接选用基于PubLayNet数据集预训练的YOLO、Mask RCNN系列模型,已经支持文本块、标题、图片、表格、列表五类常见文档元素的检测,输入PDF转换后的页面图片即可直接输出每个元素的类别和对应坐标,泛化性远高于传统规则方法
现有OpenCV代码优化建议
如果你要继续迭代当前的OpenCV实现,可以做以下调整提升效果:
- 移除
max_area=30000的过滤规则,大尺寸的表格、图片很容易被该规则误筛,可改为过滤面积极小的噪点轮廓 - 增加元素区分逻辑:通过霍夫线变换识别表格的横竖交叉线特征,先定位表格区域,排除边框干扰后再提取文本、图片的轮廓
- 文本块合并逻辑优化:不要仅使用固定的
merge_margin判断,可结合行高、左右对齐特征判断相邻轮廓是否属于同一文本块,适配不同字号的排版
参考示例
目标检测示例

文档示例图

你当前使用的OpenCV实现代码
import cv2 import numpy as np # tuplify def tup(point): return (point[0], point[1]) # returns true if the two boxes overlap def overlap(source, target): # unpack points tl1, br1 = source tl2, br2 = target # checks if (tl1[0] >= br2[0] or tl2[0] >= br1[0]): return False if (tl1[1] >= br2[1] or tl2[1] >= br1[1]): return False return True # returns all overlapping boxes def getAllOverlaps(boxes, bounds, index): overlaps = [] for a in range(len(boxes)): if a != index: if overlap(bounds, boxes[a]): overlaps.append(a) return overlaps img = cv2.imread("test.png") orig = np.copy(img) blue, green, red = cv2.split(img) def medianCanny(img, thresh1, thresh2): median = np.median(img) img = cv2.Canny(img, int(thresh1 * median), int(thresh2 * median)) return img blue_edges = medianCanny(blue, 0, 1) green_edges = medianCanny(green, 0, 1) red_edges = medianCanny(red, 0, 1) edges = blue_edges | green_edges | red_edges # I'm using OpenCV 3.4. This returns (contours, hierarchy) in OpenCV 2 and 4 _, contours,hierarchy = cv2.findContours(edges, cv2.RETR_EXTERNAL ,cv2.CHAIN_APPROX_SIMPLE) # go through the contours and save the box edges boxes = [] # each element is [[top-left], [bottom-right]] hierarchy = hierarchy[0] for component in zip(contours, hierarchy): currentContour = component[0] currentHierarchy = component[1] x,y,w,h = cv2.boundingRect(currentContour) if currentHierarchy[3] < 0: cv2.rectangle(img,(x,y),(x+w,y+h),(0,255,0),1) boxes.append([[x,y], [x+w, y+h]]) # filter out excessively large boxes filtered = [] max_area = 30000 for box in boxes: w = box[1][0] - box[0][0] h = box[1][1] - box[0][1] if w*h < max_area: filtered.append(box) boxes = filtered # go through the boxes and start merging merge_margin = 15 # this is gonna take a long time finished = False highlight = [[0,0], [1,1]] points = [[[0,0]]] while not finished: # set end con finished = True # check progress print("Len Boxes: " + str(len(boxes))) # draw boxes # comment this section out to run faster copy = np.copy(orig) for box in boxes: cv2.rectangle(copy, tup(box[0]), tup(box[1]), (0,200,0), 1) cv2.rectangle(copy, tup(highlight[0]), tup(highlight[1]), (0,0,255), 2) for point in points: point = point[0] cv2.circle(copy, tup(point), 4, (255,0,0), -1) cv2.imshow("Copy", copy) key = cv2.waitKey(1) if key == ord('q'): break # loop through boxes index = len(boxes) - 1 while index >= 0: # grab current box curr = boxes[index] # add margin tl = curr[0][:] br = curr[1][:] tl[0] -= merge_margin tl[1] -= merge_margin br[0] += merge_margin br[1] += merge_margin # get matching boxes overlaps = getAllOverlaps(boxes, [tl, br], index) # check if empty if len(overlaps) > 0: # combine boxes # convert to a contour con = [] overlaps.append(index) for ind in overlaps: tl, br = boxes[ind] con.append([tl]) con.append([br]) con = np.array(con) # get bounding rect x,y,w,h = cv2.boundingRect(con) # stop growing w -= 1 h -= 1 merged = [[x,y], [x+w, y+h]] # highlights highlight = merged[:] points = con # remove boxes from list overlaps.sort(reverse = True) for ind in overlaps: del boxes[ind] boxes.append(merged) # set flag finished = False break # increment index -= 1 cv2.destroyAllWindows() # show final copy = np.copy(orig) for box in boxes: cv2.rectangle(copy, tup(box[0]), tup(box[1]), (0,200,0), 1) cv2.imshow("Final", copy) cv2.waitKey(0)
内容的提问来源于stack exchange,提问作者Nikita Kit
相关产品推荐
相关产品推荐

