You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python检测PDF中的绘图并获取各图形独立Bounding Box

使用Python提取PDF中目标图形的Bounding Box(避免误识别短横线)

1. 先搞懂fitz.get_drawings()的返回逻辑

fitz_page.get_drawings()返回的每个元素都是PDF里的单个绘图路径——不管是复杂图形还是选项里的短横线,都会被当成独立的绘图项。每个绘图项自带bbox字段(格式为[x0, y0, x1, y1],对应左下和右上坐标),但需要我们手动过滤掉不需要的短横线。

2. 设计过滤规则

要区分目标图形和短横线,可从三个维度入手,按需组合:

  • 尺寸过滤:短横线的宽度/高度远小于目标图形,比如设置最小宽度>40、最小高度>25(根据你的PDF调整数值)
  • 位置过滤:目标图形在问题下方,先定位问题文本的位置,只提取该区域下方的图形
  • 路径复杂度过滤:短横线一般是2个点的直线,对应的draw["items"]列表长度很短;而复杂图形的路径点数多,items长度明显更大

3. 基础过滤代码示例

import fitz  # PyMuPDF

def get_target_graph_bboxes(pdf_path, page_idx=0):
    doc = fitz.open(pdf_path)
    page = doc[page_idx]
    drawings = page.get_drawings()
    
    target_bboxes = []
    # 遍历所有绘图项,过滤掉短横线
    for draw in drawings:
        bbox = draw["bbox"]
        width = bbox[2] - bbox[0]
        height = bbox[3] - bbox[1]
        # 这里的阈值根据你的PDF实际情况调整
        if width > 40 and height > 25 and len(draw["items"]) > 3:
            target_bboxes.append(bbox)
    
    doc.close()
    return target_bboxes

# 调用示例
pdf_file = "your_question_pdf.pdf"
graph_bboxes = get_target_graph_bboxes(pdf_file)
for i, bbox in enumerate(graph_bboxes):
    print(f"目标图形{i+1}的Bounding Box: {bbox}")

4. 结合文本定位的精准过滤

如果单纯尺寸过滤不够准,可以先定位问题文本的位置,只提取问题下方的图形:

def locate_question_area(page, question_keyword):
    # 搜索问题关键词,获取其位置范围
    question_matches = page.search_for(question_keyword)
    if question_matches:
        # 取第一个匹配的文本区域,以其底部y坐标作为分界
        return question_matches[0][3]
    return None

def get_target_graph_bboxes_with_text(pdf_path, page_idx=0, question_keyword="问题"):
    doc = fitz.open(pdf_path)
    page = doc[page_idx]
    
    question_bottom_y = locate_question_area(page, question_keyword)
    if not question_bottom_y:
        print("未定位到问题区域")
        doc.close()
        return []
    
    drawings = page.get_drawings()
    target_bboxes = []
    
    for draw in drawings:
        bbox = draw["bbox"]
        # 只处理问题区域下方的图形
        if bbox[1] > question_bottom_y:
            width = bbox[2] - bbox[0]
            height = bbox[3] - bbox[1]
            if width > 40 and height > 25 and len(draw["items"]) > 3:
                target_bboxes.append(bbox)
    
    doc.close()
    return target_bboxes

注意事项

  • 所有阈值(宽度、高度、items长度)需要根据你的PDF实际情况调整,建议先打印所有drawings的bbox和items字段,对比短横线和目标图形的差异后再设置
  • 如果目标图形是位图(不是矢量绘图),get_drawings()无法捕捉,需要用page.get_images()获取图片位置,再通过图片的bbox来提取坐标

内容的提问来源于stack exchange,提问作者Shivam Tripathi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 16:30:44