如何读取PDF中未知页数的跨页表格并转为Pandas DataFrame?
解决Tabula读取跨页PDF表格的问题
方法1:读取全部页面并合并结果
Tabula支持直接读取PDF的所有页面,无需指定具体页码,用pages='all'参数即可。读取后将所有页面的DataFrame合并,同时处理跨页后重复的表头行:
import pandas as pd from tabula import read_pdf # 读取所有页面的表格数据 dfs = read_pdf("your_target.pdf", pages='all', lattice=True) # 合并所有页面的DataFrame combined_df = pd.concat(dfs, ignore_index=True) # 移除跨页产生的重复表头行(假设表头为第一行内容) header = combined_df.iloc[0] combined_df = combined_df[combined_df != header].dropna(how='all') combined_df.columns = header combined_df = combined_df.reset_index(drop=True)
方法2:固定区域+Stream模式读取
如果表格在PDF每页的位置固定,可通过stream=True模式配合指定表格区域,提升跨页表格的识别连续性。区域坐标可通过Tabula桌面GUI工具获取:
from tabula import read_pdf # 读取所有页面的指定区域表格 dfs = read_pdf( "your_target.pdf", pages='all', stream=True, area=[25, 15, 740, 560], # 替换为你的表格实际坐标 guess=False # 关闭自动猜测,强制使用指定区域 ) combined_df = pd.concat(dfs, ignore_index=True)
方法3:启用合并区域参数
针对跨页时表格拆分在页面不同区域的情况,可开启merge_areas=True让Tabula自动合并相邻的表格区域:
from tabula import read_pdf dfs = read_pdf( "your_target.pdf", pages='all', lattice=True, merge_areas=True ) combined_df = pd.concat(dfs, ignore_index=True)
额外提示
- 先用Tabula桌面GUI工具预览表格,确认
Lattice或Stream模式的适配性,再将参数迁移到代码中; - 读取后若存在合并单元格导致的空值,可通过
fillna(method='ffill')等方法手动清理数据。
内容的提问来源于stack exchange,提问作者TFR
相关产品推荐
相关产品推荐

