You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从多表格PDF页面提取指定表格并转为DataFrame

精准提取PDF中指定表格的解决方案

方案1:用Camelot的table_areas参数指定提取区域

这是最直接的精准提取方式,通过页面坐标锁定目标表格:

  1. 获取目标表格坐标:
    • 用PDF阅读器(如Adobe Acrobat、Foxit)打开文件,查看目标表格的边界坐标(格式为x1,y1,x2,y2,x从左到右递增,y从下到上递增)
    • 或者用Camelot的可视化功能自动识别表格区域:
      import camelot
      tables = camelot.read_pdf("your_file.pdf", pages="1", flavor='stream', plot=True)
      tables.plot('contour')  # 弹出窗口显示页面所有识别到的表格区域
      
  2. 指定区域提取代码:
    将坐标传入table_areas参数,仅提取该区域内的表格:
    import camelot
    # 替换为你的目标表格实际坐标,示例格式:["x1,y1,x2,y2"]
    target_area = ["50,300,550,600"]
    tables = camelot.read_pdf("your_file.pdf", pages="1", flavor='stream', table_areas=target_area)
    df_target = tables[0].df
    

方案2:通过表格特征筛选目标表格

如果不想手动定位坐标,可以先提取页面所有表格,再通过表头、行列数等特征筛选目标:

import camelot

tables = camelot.read_pdf("your_file.pdf", pages="1", flavor='stream')
# 遍历所有表格,匹配目标特征(比如表头包含特定关键词)
for table in tables:
    df = table.df
    # 假设目标表格第一行是表头,包含"目标关键词"
    if "目标关键词" in df.iloc[0].values:
        df_target = df
        break

方案3:用Tabula工具实现区域提取

Tabula也是PDF表格提取的常用工具,支持通过区域参数精准定位:

from tabula import read_pdf

# area参数格式:[top, left, bottom, right],y坐标从上到下递增
target_area = [200, 50, 500, 550]
df_target = read_pdf("your_file.pdf", pages="1", area=target_area, stream=True)[0]

以上方法都能避免提取页面冗余内容,直接生成目标表格的DataFrame。如果需要通用化处理,可以把坐标或特征筛选逻辑做成可配置参数,适配不同PDF文件。

内容的提问来源于stack exchange,提问作者Zain Fendukly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 07:25:11