You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python提取PDF格式表格内的单元格图片

基于Python的PDF表格单元格切割实现方案

前置依赖安装

执行以下命令安装所需工具库:

  • pip install pymupdf opencv-python numpy

分场景实现步骤

场景1:可编辑原生PDF

这类PDF的表格结构可直接被解析,无需额外做图像识别定位单元格,用PyMuPDF(fitz)即可直接获取单元格坐标完成裁剪:

import fitz
import cv2
import numpy as np

# 1. 打开PDF并读取指定页
pdf_path = "你的文件路径.pdf"
doc = fitz.open(pdf_path)
page = doc[0]  # 取第一页,按需修改页码

# 2. 提取页面内所有表格
tables = page.find_tables()

# 3. 将PDF页转为高分辨率图片(dpi设为300保证手写内容清晰度)
zoom = 300 / 72
matrix = fitz.Matrix(zoom, zoom)
pix = page.get_pixmap(matrix=matrix)
img = np.frombuffer(pix.samples, dtype=np.uint8).reshape(pix.height, pix.width, pix.n)
if pix.n == 4:  # 转RGB格式
    img = cv2.cvtColor(img, cv2.COLOR_BGRA2BGR)

# 4. 遍历所有表格的所有单元格,裁剪保存
for table_idx, table in enumerate(tables):
    for row_idx, row in enumerate(table.cells):
        for col_idx, cell in enumerate(row):
            # 获取单元格原始坐标,适配缩放后的图片尺寸
            x0, y0, x1, y1 = [i * zoom for i in cell]
            # 裁剪单元格
            cell_img = img[int(y0):int(y1), int(x0):int(x1)]
            # 保存图片,命名带表格、行、列序号方便后续整理
            cv2.imwrite(f"table_{table_idx}_row_{row_idx}_col_{col_idx}.png", cell_img)

场景2:扫描件/图片类PDF

这类PDF的内容全为栅格图像,无法直接解析表格结构,需先转成图片后通过OpenCV做表格检测定位单元格:

import fitz
import cv2
import numpy as np

# 1. PDF转高分辨率图片
pdf_path = "你的扫描件路径.pdf"
doc = fitz.open(pdf_path)
page = doc[0]
zoom = 300 / 72
matrix = fitz.Matrix(zoom, zoom)
pix = page.get_pixmap(matrix=matrix)
img = np.frombuffer(pix.samples, dtype=np.uint8).reshape(pix.height, pix.width, pix.n)
if img.shape[-1] == 4:
    img = cv2.cvtColor(img, cv2.COLOR_BGRA2BGR)
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)

# 2. 阈值处理+边缘检测
thresh = cv2.adaptiveThreshold(gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)

# 3. 检测表格横竖线
horizontal_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (40, 1))
detect_horizontal = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, horizontal_kernel, iterations=2)
cnts_horizontal = cv2.findContours(detect_horizontal, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)[0]

vertical_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (1, 40))
detect_vertical = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, vertical_kernel, iterations=2)
cnts_vertical = cv2.findContours(detect_vertical, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)[0]

# 4. 合并横竖线得到表格掩码
table_mask = np.zeros_like(thresh)
cv2.drawContours(table_mask, cnts_horizontal, -1, 255, 2)
cv2.drawContours(table_mask, cnts_vertical, -1, 255, 2)

# 5. 查找单元格轮廓
cnts = cv2.findContours(table_mask, cv2.RETR_TREE, cv2.CHAIN_APPROX_SIMPLE)[0]
# 过滤非单元格的小轮廓,面积阈值可根据实际表格尺寸调整
cells = []
for c in cnts:
    area = cv2.contourArea(c)
    if 1000 < area < img.shape[0]*img.shape[1]*0.1:
        x, y, w, h = cv2.boundingRect(c)
        cells.append((x, y, x+w, y+h))

# 6. 按行、列排序后裁剪保存,排序逻辑可根据表格格式自定义
cells = sorted(cells, key=lambda x: (x[1], x[0]))
for idx, (x0, y0, x1, y1) in enumerate(cells):
    cell_img = img[y0:y1, x0:x1]
    cv2.imwrite(f"cell_{idx}.png", cell_img)

后续处理说明

裁剪得到的单元格图片可直接传入手写OCR模型完成内容识别,最终按表格的行列对应关系整理结果即可输出为JSON格式。

内容的提问来源于stack exchange,提问作者pip install

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 07:57:03