You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Tabula提取PDF表格时每页最后一行丢失的解决方法咨询

Tabula提取PDF表格时每页最后一行丢失的解决办法

针对你用Tabula提取4页PDF表格(共98行)时每页最后一行未被提取的问题,可尝试以下几种解决方案:

1. 手动指定页面提取区域

Tabula默认的页面识别区域可能未覆盖到表格最后一行,可通过area参数自定义提取范围。参数格式为[top, left, bottom, right],数值对应PDF页面的像素坐标(可通过PDF阅读器查看页面尺寸,或用Tabula的GUI工具先测试合适的区域)。

示例代码:

import tabula
# 假设页面尺寸为842x595(A4),手动调整bottom值确保覆盖最后一行
tabula.convert_into('Path/abc.pdf', 'Path/abc.csv', output_format='csv', pages='all', lattice=True, area=[0, 0, 840, 595])

如果需要用相对坐标(0-1的比例值),可加上relative_area=True:

tabula.convert_into('Path/abc.pdf', 'Path/abc.csv', output_format='csv', pages='all', lattice=True, area=[0, 0, 0.99, 1], relative_area=True)

2. 切换到Stream模式提取

lattice模式依赖表格边框识别,若最后一行的边框不清晰或存在格式问题,可尝试stream=True(基于文本位置识别表格):

import tabula
tabula.convert_into('Path/abc.pdf', 'Path/abc.csv', output_format='csv', pages='all', stream=True)

3. 分页面读取后合并导出

先通过read_pdf读取每页表格,检查是否缺失行,再合并导出为CSV,这种方式更便于调试:

import tabula
import pandas as pd

# 读取所有页面的表格
dfs = tabula.read_pdf('Path/abc.pdf', pages='all', lattice=True, multiple_tables=False)
# 合并所有DataFrame
combined_df = pd.concat(dfs)
# 导出到CSV
combined_df.to_csv('Path/abc.csv', index=False)

如果发现某页缺失行,可单独调整该页的area参数后重新读取:

# 单独读取第1页,指定自定义区域
df_page1 = tabula.read_pdf('Path/abc.pdf', pages=1, lattice=True, area=[0, 0, 840, 595])
# 读取其他页
df_other = tabula.read_pdf('Path/abc.pdf', pages='2-4', lattice=True)
# 合并
combined_df = pd.concat([df_page1[0]] + df_other)
combined_df.to_csv('Path/abc.csv', index=False)

4. 调整表格识别参数

尝试添加guess=False参数(关闭自动猜测表格区域),配合area参数强制使用指定区域:

tabula.convert_into('Path/abc.pdf', 'Path/abc.csv', output_format='csv', pages='all', lattice=True, area=[0, 0, 840, 595], guess=False)

内容的提问来源于stack exchange,提问作者Mukesh Dalmia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 20:12:33