You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python提取PDF中各元素(文本、图片、表格)的坐标

PDF文档元素坐标提取方案

这个需求完全可以实现,你当前使用传统OpenCV轮廓检测+Tesseract的方案效果不佳,核心问题是规则化的传统视觉方法无法适配复杂排版的文档,也不能对文本、图片、表格三类元素做分类识别。

现有方案存在的问题

  • 纯OpenCV轮廓检测仅能识别边界特征,无法区分轮廓对应的元素类型,很容易把文本行、表格边框、图片边缘混判,也无法自动合并属于同一文本块的分散轮廓
  • 单独使用Tesseract时如果没有做前置区域分割,会出现文本块拆分错误、坐标偏移的问题,也无法输出表格、图片的坐标信息

可落地的实现方案

轻量原生PDF适配方案(无需训练,仅适用于可复制的非扫描PDF)

不需要将PDF转成图片用CV处理,直接调用PDF解析工具即可精准获取所有元素坐标:

  • 文本块:使用PyMuPDF库的page.get_text("blocks")接口,返回结果直接包含每个文本块的左上角、右下角坐标以及对应文本内容
  • 图片:使用PyMuPDF库的page.get_images()接口,可直接获取每个插入图片的坐标、大小与原始数据
  • 表格:配合camelot或tabula-py工具提取表格内容的同时,即可获取对应表格的页面坐标,识别准确率远高于OpenCV边框检测

深度学习通用方案(适配扫描版PDF、复杂排版文档)

你示例中用到的目标检测方案是当前通用场景下准确率最高的方案,不需要从零开发训练,直接使用成熟预训练模型即可:

  • 可直接选用基于PubLayNet数据集预训练的YOLO、Mask RCNN系列模型,已经支持文本块、标题、图片、表格、列表五类常见文档元素的检测,输入PDF转换后的页面图片即可直接输出每个元素的类别和对应坐标,泛化性远高于传统规则方法

现有OpenCV代码优化建议

如果你要继续迭代当前的OpenCV实现,可以做以下调整提升效果:

  • 移除max_area=30000的过滤规则,大尺寸的表格、图片很容易被该规则误筛,可改为过滤面积极小的噪点轮廓
  • 增加元素区分逻辑:通过霍夫线变换识别表格的横竖交叉线特征,先定位表格区域,排除边框干扰后再提取文本、图片的轮廓
  • 文本块合并逻辑优化:不要仅使用固定的merge_margin判断,可结合行高、左右对齐特征判断相邻轮廓是否属于同一文本块,适配不同字号的排版

参考示例

目标检测示例

目标检测示例图

文档示例图

文档示例图

你当前使用的OpenCV实现代码

import cv2
import numpy as np

# tuplify
def tup(point):
    return (point[0], point[1])

# returns true if the two boxes overlap
def overlap(source, target):
    # unpack points
    tl1, br1 = source
    tl2, br2 = target

    # checks
    if (tl1[0] >= br2[0] or tl2[0] >= br1[0]):
        return False
    if (tl1[1] >= br2[1] or tl2[1] >= br1[1]):
        return False
    return True

# returns all overlapping boxes
def getAllOverlaps(boxes, bounds, index):
    overlaps = []
    for a in range(len(boxes)):
        if a != index:
            if overlap(bounds, boxes[a]):
                overlaps.append(a)
    return overlaps

img = cv2.imread("test.png")
orig = np.copy(img)
blue, green, red = cv2.split(img)

def medianCanny(img, thresh1, thresh2):
    median = np.median(img)
    img = cv2.Canny(img, int(thresh1 * median), int(thresh2 * median))
    return img

blue_edges = medianCanny(blue, 0, 1)
green_edges = medianCanny(green, 0, 1)
red_edges = medianCanny(red, 0, 1)

edges = blue_edges | green_edges | red_edges

# I'm using OpenCV 3.4. This returns (contours, hierarchy) in OpenCV 2 and 4
_, contours,hierarchy = cv2.findContours(edges, cv2.RETR_EXTERNAL ,cv2.CHAIN_APPROX_SIMPLE)

# go through the contours and save the box edges
boxes = [] # each element is [[top-left], [bottom-right]]
hierarchy = hierarchy[0]
for component in zip(contours, hierarchy):
    currentContour = component[0]
    currentHierarchy = component[1]
    x,y,w,h = cv2.boundingRect(currentContour)
    if currentHierarchy[3] < 0:
        cv2.rectangle(img,(x,y),(x+w,y+h),(0,255,0),1)
        boxes.append([[x,y], [x+w, y+h]])

# filter out excessively large boxes
filtered = []
max_area = 30000
for box in boxes:
    w = box[1][0] - box[0][0]
    h = box[1][1] - box[0][1]
    if w*h < max_area:
        filtered.append(box)
boxes = filtered

# go through the boxes and start merging
merge_margin = 15

# this is gonna take a long time
finished = False
highlight = [[0,0], [1,1]]
points = [[[0,0]]]
while not finished:
    # set end con
    finished = True

    # check progress
    print("Len Boxes: " + str(len(boxes)))

    # draw boxes # comment this section out to run faster
    copy = np.copy(orig)
    for box in boxes:
        cv2.rectangle(copy, tup(box[0]), tup(box[1]), (0,200,0), 1)
    cv2.rectangle(copy, tup(highlight[0]), tup(highlight[1]), (0,0,255), 2)
    for point in points:
        point = point[0]
        cv2.circle(copy, tup(point), 4, (255,0,0), -1)
    cv2.imshow("Copy", copy)
    key = cv2.waitKey(1)
    if key == ord('q'):
        break

    # loop through boxes
    index = len(boxes) - 1
    while index >= 0:
        # grab current box
        curr = boxes[index]

        # add margin
        tl = curr[0][:]
        br = curr[1][:]
        tl[0] -= merge_margin
        tl[1] -= merge_margin
        br[0] += merge_margin
        br[1] += merge_margin

        # get matching boxes
        overlaps = getAllOverlaps(boxes, [tl, br], index)
        
        # check if empty
        if len(overlaps) > 0:
            # combine boxes
            # convert to a contour
            con = []
            overlaps.append(index)
            for ind in overlaps:
                tl, br = boxes[ind]
                con.append([tl])
                con.append([br])
            con = np.array(con)

            # get bounding rect
            x,y,w,h = cv2.boundingRect(con)

            # stop growing
            w -= 1
            h -= 1
            merged = [[x,y], [x+w, y+h]]

            # highlights
            highlight = merged[:]
            points = con

            # remove boxes from list
            overlaps.sort(reverse = True)
            for ind in overlaps:
                del boxes[ind]
            boxes.append(merged)

            # set flag
            finished = False
            break

        # increment
        index -= 1
cv2.destroyAllWindows()

# show final
copy = np.copy(orig)
for box in boxes:
    cv2.rectangle(copy, tup(box[0]), tup(box[1]), (0,200,0), 1)
cv2.imshow("Final", copy)
cv2.waitKey(0)

内容的提问来源于stack exchange,提问作者Nikita Kit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 08:21:00