You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Colab使用pdfminer提取PDF文本坐标并绘制字符矩形框求助

PDF文本坐标识别+矩形框绘制Colab实现方案

首先执行以下命令安装依赖:

!pip install pdfminer.six pymupdf

你之前的代码存在三个核心问题:

  • 文件打开顺序错误,open(pathh, 'rb')执行时pathh还未赋值,会直接报错
  • pdfminer返回的坐标原点是页面左下角,常见绘图库的坐标原点是左上角,未做坐标转换就绘制会出现位置偏移
  • 只遍历到了文本块LTTextBox层级,要获取单个字符的坐标需要向下遍历到LTChar对象

完整调整后代码

from pdfminer.layout import LAParams, LTTextBox, LTTextLine, LTChar
from pdfminer.pdfpage import PDFPage
from pdfminer.pdfinterp import PDFResourceManager, PDFPageInterpreter
from pdfminer.converter import PDFPageAggregator
import fitz
from google.colab import drive
from PIL import Image

# 挂载谷歌云盘
drive.mount('/content/drive')
# 配置文件路径
pathh = "/content/drive/MyDrive/ICAO.pdf"
output_path = "/content/ICAO_with_boxes.pdf"

# 第一步:用pdfminer提取所有字符/文本块的坐标
rsrcmgr = PDFResourceManager()
laparams = LAParams()
device = PDFPageAggregator(rsrcmgr, laparams=laparams)
interpreter = PDFPageInterpreter(rsrcmgr, device)
fp = open(pathh, 'rb')
pages = list(PDFPage.get_pages(fp))
# 存储每页的文本坐标信息
page_text_coords = []

for page in pages:
    interpreter.process_page(page)
    layout = device.get_result()
    page_width = page.mediabox[2]
    page_height = page.mediabox[3]
    current_page_coords = []
    for lobj in layout:
        if isinstance(lobj, LTTextBox):
            # 遍历到单个字符层级,要文本块级坐标的话可以直接取lobj.bbox
            for line in lobj:
                if isinstance(line, LTTextLine):
                    for char in line:
                        if isinstance(char, LTChar):
                            x0, y0, x1, y1 = char.bbox
                            # 坐标转换:把左下角原点转成左上角原点
                            y0_new = page_height - y1
                            y1_new = page_height - y0
                            current_page_coords.append({
                                "x0": x0, "y0": y0_new,
                                "x1": x1, "y1": y1_new,
                                "text": char.get_text()
                            })
    page_text_coords.append({
        "width": page_width,
        "height": page_height,
        "coords": current_page_coords
    })
fp.close()

# 第二步:用PyMuPDF绘制矩形框
doc = fitz.open(pathh)
for page_idx in range(len(doc)):
    page = doc[page_idx]
    coords_info = page_text_coords[page_idx]
    for item in coords_info["coords"]:
        # 绘制红色矩形框,线宽0.5
        rect = fitz.Rect(item["x0"], item["y0"], item["x1"], item["y1"])
        page.draw_rect(rect, color=(1,0,0), width=0.5)
# 保存带框的PDF
doc.save(output_path)
doc.close()

# 可选:在Colab中直接展示第一页效果
doc = fitz.open(output_path)
page = doc[0]
pix = page.get_pixmap()
img = Image.frombytes("RGB", [pix.width, pix.height], pix.samples)
display(img)

效果调整说明

  • 要绘制文本块级别的框,只需把LTChar层级的遍历逻辑去掉,直接处理LTTextBox的bbox即可
  • 矩形框的颜色、线宽可以修改draw_rect方法的color、width参数调整
  • 不需要单字符坐标的话可以大幅简化遍历逻辑,运行速度也会更快

内容的提问来源于stack exchange,提问作者engineering Baba

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 04:24:01