使用Camelot提取PDF表格遇DeprecationError,求正确操作步骤
解决Camelot提取PDF表格的错误及正确步骤
一、解决PyPDF2版本冲突问题
你遇到的DeprecationError是因为Camelot当前版本未适配PyPDF2 3.0.0+的接口变更(旧的PdfFileReader被移除,替换为PdfReader),需要安装兼容的PyPDF2版本:
- 卸载当前PyPDF2:
pip uninstall -y PyPDF2 - 安装2.x稳定版:
pip install PyPDF2==2.12.1
如果在Google Colab中操作,直接在代码单元格运行上述命令即可。
二、Camelot提取PDF表格的完整步骤
1. 安装依赖
确保配齐核心工具:
- 安装Camelot:
pip install camelot-py[cv](带[cv]可支持可视化调试功能) - GhostScript配置:
- Windows:下载对应版本安装后,将GhostScript的
bin目录添加到系统环境变量PATH - Linux/macOS:通过包管理器安装,比如Ubuntu用
apt install ghostscript,macOS用brew install ghostscript - Colab无需手动安装,系统已预装GhostScript
- Windows:下载对应版本安装后,将GhostScript的
2. 基础提取代码
import camelot # 读取PDF所有页面的表格 tables = camelot.read_pdf(r"F:\testing\sbi_9.pdf", pages="all") # 查看提取到的表格总数 print(f"提取到{len(tables)}个表格") # 打印第一个表格的内容(以DataFrame格式输出) print(tables[0].df) # 将所有表格导出为CSV文件 tables.export("extracted_tables.csv", f="csv", compress=False)
3. 提取效果优化(可选)
如果表格识别不准确,可调整参数:
flavor参数:lattice适合带明确边框的表格,stream适合无边框表格,例如camelot.read_pdf("file.pdf", flavor="stream", pages="all")edge_tol参数:调整边框识别的容忍度,比如edge_tol=500- 可视化识别区域:
camelot.plot(tables[0], kind="contour").show(),帮助定位识别偏差点
内容的提问来源于stack exchange,提问作者Raj
相关产品推荐
相关产品推荐

