You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Tabula提取PDF表格时同步获取表格外的对应标题?

解决Tabula提取表格同时匹配标题的方案

这里提供两种实用方法,既能保留Tabula的表格提取精度,又能准确获取表格上方的对应标题:

方法一:结合PyMuPDF(fitz)匹配表格位置与标题

利用PyMuPDF提取页面文本的位置信息,和Tabula识别的表格坐标做匹配,筛选出表格上方的对应文本作为标题,是最通用的方案。

步骤:

  1. 安装依赖包
pip install tabula-py pymupdf
  1. 示例代码
import tabula.io as tb
import fitz  # PyMuPDF

file_path = "你的PDF文件路径.pdf"
target_pages = "1"  # 支持多页,格式如"1-5"、"2,4"

# 1. Tabula提取表格及坐标信息(用JSON格式获取位置)
tables = tb.read_pdf(
    file_path,
    pages=target_pages,
    multiple_tables=True,
    output_format="json"
)

# 2. PyMuPDF读取PDF页面
doc = fitz.open(file_path)

for page_idx in range(len(doc)):
    current_page_num = page_idx + 1
    # 跳过非目标页面
    if str(current_page_num) not in target_pages.replace(",", "-").split("-"):
        continue
    
    page = doc[page_idx]
    # 获取页面所有文本块,每个块包含:(x左, y上, x右, y下, 文本内容, 块编号, 类型)
    text_blocks = page.get_text("blocks")
    
    # 遍历当前页的每个表格,匹配标题
    for table in tables:
        if table["page"] != current_page_num:
            continue
        
        # 获取表格的坐标边界
        table_top = table["top"]
        table_left = table["left"]
        table_right = table["right"]
        
        # 筛选表格上方的文本块:文本块底部在表格顶部上方,且水平范围与表格有重叠
        title_candidates = []
        for block in text_blocks:
            blk_x0, blk_y0, blk_x1, blk_y1, blk_text = block[:5]
            cleaned_text = blk_text.strip()
            if not cleaned_text:
                continue
            # 判断文本块是否在表格上方且水平重叠
            if blk_y1 < table_top and not (blk_x1 < table_left or blk_x0 > table_right):
                title_candidates.append( (blk_y1, cleaned_text) )
        
        # 取距离表格最近的文本块作为标题(按y坐标从大到小排序)
        if title_candidates:
            title_candidates.sort(reverse=True, key=lambda x: x[0])
            table_title = title_candidates[0][1]
            print(f"【表格标题】:{table_title}")
        else:
            print("【表格标题】:未找到对应标题")
        
        # 将表格转为DataFrame(按需使用)
        table_df = tb.convert_into(table, output_format="dataframe", pages=current_page_num)[0]
        print("【表格内容】:")
        print(table_df)
        print("=" * 60)

doc.close()

注意事项:

  • 如果标题是多行文本,可以调整逻辑,将连续的上方文本块合并(比如判断文本块之间的垂直距离是否小于阈值)
  • 若PDF存在复杂排版(如标题和表格水平偏移较大),可微调水平重叠的判断条件

方法二:固定位置范围提取标题(适用于排版规整的PDF)

如果你的PDF排版固定,标题和表格的相对位置(比如标题在表格上方固定高度范围内),可以直接用Tabula指定区域分别提取标题和表格:

import tabula.io as tb

file_path = "你的PDF文件路径.pdf"

# 提取标题:指定标题所在的页面区域(坐标需自己调整,可先用Tabula的GUI工具查看坐标)
title = tb.read_pdf(
    file_path,
    pages="1",
    area=[[20, 20, 80, 500]],  # [top, left, bottom, right]
    output_format="plain"
)[0].strip()

# 提取表格:指定表格所在的页面区域
table_df = tb.read_pdf(
    file_path,
    pages="1",
    area=[[100, 20, 500, 500]],
    output_format="dataframe"
)[0]

print(f"标题:{title}")
print("表格:")
print(table_df)

注意事项:

  • 坐标需要根据实际PDF调整,可使用tabula.gui()打开Tabula的GUI工具查看区域坐标

内容的提问来源于stack exchange,提问作者user15410844

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 00:30:47