You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Camelot提取PDF表格遇DeprecationError,求正确操作步骤

解决Camelot提取PDF表格的错误及正确步骤

一、解决PyPDF2版本冲突问题

你遇到的DeprecationError是因为Camelot当前版本未适配PyPDF2 3.0.0+的接口变更(旧的PdfFileReader被移除,替换为PdfReader),需要安装兼容的PyPDF2版本:

  • 卸载当前PyPDF2:pip uninstall -y PyPDF2
  • 安装2.x稳定版:pip install PyPDF2==2.12.1

如果在Google Colab中操作,直接在代码单元格运行上述命令即可。

二、Camelot提取PDF表格的完整步骤

1. 安装依赖

确保配齐核心工具:

  • 安装Camelot:pip install camelot-py[cv](带[cv]可支持可视化调试功能)
  • GhostScript配置:
    • Windows:下载对应版本安装后,将GhostScript的bin目录添加到系统环境变量PATH
    • Linux/macOS:通过包管理器安装,比如Ubuntu用apt install ghostscript,macOS用brew install ghostscript
    • Colab无需手动安装,系统已预装GhostScript

2. 基础提取代码

import camelot

# 读取PDF所有页面的表格
tables = camelot.read_pdf(r"F:\testing\sbi_9.pdf", pages="all")

# 查看提取到的表格总数
print(f"提取到{len(tables)}个表格")

# 打印第一个表格的内容(以DataFrame格式输出)
print(tables[0].df)

# 将所有表格导出为CSV文件
tables.export("extracted_tables.csv", f="csv", compress=False)

3. 提取效果优化(可选)

如果表格识别不准确,可调整参数:

  • flavor参数:lattice适合带明确边框的表格,stream适合无边框表格,例如camelot.read_pdf("file.pdf", flavor="stream", pages="all")
  • edge_tol参数:调整边框识别的容忍度,比如edge_tol=500
  • 可视化识别区域:camelot.plot(tables[0], kind="contour").show(),帮助定位识别偏差点

内容的提问来源于stack exchange,提问作者Raj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 17:32:06