You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Tabula-py提取PDF中的表格及其前置标题?

提取PDF表格及对应前置标题的实现方法

核心思路

Tabula本身专注于表格提取,要同时获取前置标题,需要结合表格位置定位和页面文本提取工具配合实现,这里用PyMuPDF(fitz)来提取标题文本,步骤如下:

具体实现

  1. 先安装依赖包:

    pip install tabula-py pymupdf
    
  2. 示例代码:

    import tabula
    import fitz
    
    # 替换为你的PDF文件路径
    pdf_file = "target.pdf"
    
    # 提取PDF中所有表格的位置和结构信息(用json格式输出方便获取坐标)
    table_info_list = tabula.read_pdf(
        pdf_file,
        pages="all",
        lattice=True,
        guess=False,
        output_format="json"
    )
    
    # 打开PDF文件用于文本提取
    pdf_doc = fitz.open(pdf_file)
    
    for table_info in table_info_list:
        # 转换页码:Tabula的页码从1开始,PyMuPDF从0开始
        page_idx = table_info["page"] - 1
        current_page = pdf_doc[page_idx]
    
        # 获取表格的顶部坐标
        table_top_y = table_info["top"]
        # 定义标题区域:表格上方50像素范围内(可根据实际PDF调整高度)
        title_rect = fitz.Rect(
            0,  # 左边界
            table_top_y - 50,  # 上边界(表格上方50px)
            current_page.rect.width,  # 右边界
            table_top_y  # 下边界(表格顶部)
        )
        # 提取标题文本并去除多余空白
        table_title = current_page.get_textbox(title_rect).strip()
    
        # 根据表格坐标精准提取表格内容
        table_df = tabula.read_pdf(
            pdf_file,
            pages=page_idx + 1,
            area=[table_top_y, table_info["left"], table_info["bottom"], table_info["right"]],
            lattice=True
        )[0]
    
        # 输出结果
        print(f"=== 标题: {table_title} ===")
        print(table_df)
        print("\n")
    

适配调整

  • 如果标题和表格间距较大,可修改table_top_y - 50中的数值,扩大标题提取区域;
  • 部分PDF标题可能是多行,可通过检查文本块的垂直位置,自动找到表格上方最近的连续文本块作为标题;
  • 若PDF是流式而非格子型表格,可将lattice=True改为stream=True。

内容的提问来源于stack exchange,提问作者Joel Gissing

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 07:39:59