You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Python-camelot获取的表格坐标提取对应标题区域坐标?

基于Camelot表格坐标获取PDF标题区域的Python方案

刚好之前做过类似的需求,结合Camelot和PyMuPDF(fitz)就能搞定这个问题。首先得明确几个关键点:Camelot返回的table._bbox格式是(x_left, y_bottom, x_right, y_top)——因为PDF的坐标系是以左下角为原点,y轴向上递增的,这个坐标逻辑一定要搞对,不然区域计算全错。

下面分步骤给你讲具体实现:

1. 先装依赖

需要用到Camelot提取表格,PyMuPDF处理PDF文本和区域:

pip install camelot-py[cv] pymupdf

2. 核心逻辑:计算标题候选区域

标题可能在表格的上方、左侧或下方,我们可以基于表格的bbox,加上自定义的边距和标题尺寸,计算出这三个方向的候选区域。同时要过滤掉超出页面范围的区域(比如表格在页面最左边,左侧标题区域就会超出页面,直接忽略)。

import camelot
import fitz  # PyMuPDF

def calculate_title_regions(table_bbox, page_rect, margin=15, title_height=50, title_width=80):
    """
    计算表格的标题候选区域坐标
    :param table_bbox: Camelot返回的表格bbox,格式(x1, y1, x2, y2),x1左,y1下,x2右,y2上
    :param page_rect: PyMuPDF页面的rect对象,包含页面宽高
    :param margin: 标题和表格之间的边距
    :param title_height: 标题区域的高度(上下方向)
    :param title_width: 标题区域的宽度(左右方向)
    :return: 有效标题区域字典,key是位置(top/left/bottom),value是区域坐标
    """
    x1, y1, x2, y2 = table_bbox
    page_width = page_rect.width
    page_height = page_rect.height
    regions = {}
    
    # 上方标题区域:和表格左右对齐,在表格上方margin距离处,高度title_height
    top_region = (x1, y2 + margin, x2, y2 + margin + title_height)
    # 检查是否在页面范围内
    if top_region[3] <= page_height:
        regions["top"] = top_region
    
    # 左侧标题区域:和表格上下对齐,在表格左侧margin距离处,宽度title_width
    left_region = (x1 - margin - title_width, y1, x1 - margin, y2)
    if left_region[0] >= 0:
        regions["left"] = left_region
    
    # 下方标题区域:和表格左右对齐,在表格下方margin距离处,高度title_height
    bottom_region = (x1, y1 - margin - title_height, x2, y1 - margin)
    if bottom_region[1] >= 0:
        regions["bottom"] = bottom_region
    
    return regions

3. 完整流程:提取表格+计算标题区域+验证文本

接下来把整个流程串起来,遍历每个表格,计算标题区域,还可以提取区域内的文本,甚至通过字体特征判断是不是标题(比如加粗、字号更大)。

def is_title_span(span):
    """判断文本片段是否是标题:字号大于12,加粗,非空"""
    is_bold = (span["flags"] & 1) != 0  # PyMuPDF中flags=1表示加粗
    return span["size"] > 12 and is_bold and len(span["text"].strip()) > 0

# 加载PDF
pdf_path = "your_target.pdf"
tables = camelot.read_pdf(pdf_path, pages="all")  # pages参数按需调整
doc = fitz.open(pdf_path)

# 遍历每个表格处理
for table_idx, table in enumerate(tables, 1):
    page_num = table.page  # Camelot的页码从1开始
    page = doc[page_num - 1]  # PyMuPDF的页码从0开始
    table_bbox = table._bbox
    page_rect = page.rect
    
    # 计算标题区域
    title_regions = calculate_title_regions(table_bbox, page_rect)
    
    print(f"\n=== 表格{table_idx}(页码{page_num})===")
    print(f"表格bbox: {table_bbox}")
    
    # 遍历每个候选区域,提取并验证标题
    for pos, region in title_regions.items():
        print(f"\n{pos}方向标题区域: {region}")
        # 获取区域内的文本块(带字体信息)
        text_dict = page.get_text("dict", clip=region)
        title_content = ""
        for block in text_dict.get("blocks", []):
            if block["type"] == 0:  # 只处理文本块
                for line in block["lines"]:
                    for span in line["spans"]:
                        if is_title_span(span):
                            title_content += span["text"] + " "
        
        if title_content:
            print(f"识别到标题: {title_content.strip()}")
        else:
            # 如果没有符合字体特征的,就取区域内所有文本备用
            raw_text = page.get_textbox(region).strip()
            print(f"未识别到符合特征的标题,区域原始文本: {raw_text}")

4. 调优建议

  • 边距和尺寸参数:根据你的PDF排版调整margin、title_height、title_width——如果标题和表格靠得近,margin设小一点(比如10);如果是多行标题,title_height调大(比如70)。
  • 页面尺寸:代码里用了页面实际尺寸来过滤无效区域,避免出现负数坐标或者超出页面的情况,比固定A4尺寸更靠谱。
  • 标题判断逻辑:如果你的标题没有加粗或者字号差异不大,可以调整is_title_span的条件,比如只看文本长度,或者结合关键词(比如包含"表"、"Table"等)。

内容的提问来源于stack exchange,提问作者jessy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:03:17