如何使用Tabula-py提取PDF中的表格及其前置标题?
提取PDF表格及对应前置标题的实现方法
核心思路
Tabula本身专注于表格提取,要同时获取前置标题,需要结合表格位置定位和页面文本提取工具配合实现,这里用PyMuPDF(fitz)来提取标题文本,步骤如下:
具体实现
先安装依赖包:
pip install tabula-py pymupdf示例代码:
import tabula import fitz # 替换为你的PDF文件路径 pdf_file = "target.pdf" # 提取PDF中所有表格的位置和结构信息(用json格式输出方便获取坐标) table_info_list = tabula.read_pdf( pdf_file, pages="all", lattice=True, guess=False, output_format="json" ) # 打开PDF文件用于文本提取 pdf_doc = fitz.open(pdf_file) for table_info in table_info_list: # 转换页码:Tabula的页码从1开始,PyMuPDF从0开始 page_idx = table_info["page"] - 1 current_page = pdf_doc[page_idx] # 获取表格的顶部坐标 table_top_y = table_info["top"] # 定义标题区域:表格上方50像素范围内(可根据实际PDF调整高度) title_rect = fitz.Rect( 0, # 左边界 table_top_y - 50, # 上边界(表格上方50px) current_page.rect.width, # 右边界 table_top_y # 下边界(表格顶部) ) # 提取标题文本并去除多余空白 table_title = current_page.get_textbox(title_rect).strip() # 根据表格坐标精准提取表格内容 table_df = tabula.read_pdf( pdf_file, pages=page_idx + 1, area=[table_top_y, table_info["left"], table_info["bottom"], table_info["right"]], lattice=True )[0] # 输出结果 print(f"=== 标题: {table_title} ===") print(table_df) print("\n")
适配调整
- 如果标题和表格间距较大,可修改
table_top_y - 50中的数值,扩大标题提取区域; - 部分PDF标题可能是多行,可通过检查文本块的垂直位置,自动找到表格上方最近的连续文本块作为标题;
- 若PDF是流式而非格子型表格,可将
lattice=True改为stream=True。
内容的提问来源于stack exchange,提问作者Joel Gissing
相关产品推荐
相关产品推荐

