You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何读取PDF中未知页数的跨页表格并转为Pandas DataFrame?

解决Tabula读取跨页PDF表格的问题

方法1:读取全部页面并合并结果

Tabula支持直接读取PDF的所有页面,无需指定具体页码,用pages='all'参数即可。读取后将所有页面的DataFrame合并,同时处理跨页后重复的表头行:

import pandas as pd
from tabula import read_pdf

# 读取所有页面的表格数据
dfs = read_pdf("your_target.pdf", pages='all', lattice=True)

# 合并所有页面的DataFrame
combined_df = pd.concat(dfs, ignore_index=True)

# 移除跨页产生的重复表头行(假设表头为第一行内容)
header = combined_df.iloc[0]
combined_df = combined_df[combined_df != header].dropna(how='all')
combined_df.columns = header
combined_df = combined_df.reset_index(drop=True)

方法2:固定区域+Stream模式读取

如果表格在PDF每页的位置固定,可通过stream=True模式配合指定表格区域,提升跨页表格的识别连续性。区域坐标可通过Tabula桌面GUI工具获取:

from tabula import read_pdf

# 读取所有页面的指定区域表格
dfs = read_pdf(
    "your_target.pdf",
    pages='all',
    stream=True,
    area=[25, 15, 740, 560],  # 替换为你的表格实际坐标
    guess=False  # 关闭自动猜测,强制使用指定区域
)
combined_df = pd.concat(dfs, ignore_index=True)

方法3:启用合并区域参数

针对跨页时表格拆分在页面不同区域的情况,可开启merge_areas=True让Tabula自动合并相邻的表格区域:

from tabula import read_pdf

dfs = read_pdf(
    "your_target.pdf",
    pages='all',
    lattice=True,
    merge_areas=True
)
combined_df = pd.concat(dfs, ignore_index=True)

额外提示

  • 先用Tabula桌面GUI工具预览表格,确认Lattice或Stream模式的适配性,再将参数迁移到代码中;
  • 读取后若存在合并单元格导致的空值,可通过fillna(method='ffill')等方法手动清理数据。

内容的提问来源于stack exchange,提问作者TFR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 13:24:59