You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python提取指定PDF文本和表格并存为.csv文件,解决PyPDF2、Camelot报错

问题解决及需求实现方案

报错修复

camelot报错OSError: Ghostscript is not installed解决

根据你使用的操作系统执行对应操作:

  • Windows:下载对应系统位数的Ghostscript官方安装包,安装完成后将安装目录下的bin文件夹路径(例:C:\Program Files\gs\gs10.03.0\bin)添加到系统环境变量PATH,重启IDE/命令行后生效
  • Mac:执行命令brew install ghostscript完成安装
  • Linux:执行命令sudo apt install ghostscript完成安装

PyPDF2提取空白原因说明

PyPDF2仅支持提取原生可复制文本类PDF的内容,空白输出是因为目标PDF为扫描件、加密或者字体编码特殊,无法直接用PyPDF2提取结构化内容,更适合用camelot、pdfplumber这类专门的表格提取工具。

完整功能实现代码

首先执行命令安装所需依赖:
pip install camelot-py[cv] pandas
适配需求的代码如下,字段索引需根据你实际PDF的表格结构调整:

import camelot
import pandas as pd

# 读取PDF所有页的表格,有边框表格用flavor='lattice',无边框用flavor='stream'
tables = camelot.read_pdf(r"E:\user\programs\28_oct_bank_statement\demo.pdf", pages='all', flavor='lattice')

# 提取第一张表格的账号、支行名称字段
first_table = tables[0].df
# 可先打印first_table查看内容,调整iloc的行列索引匹配实际字段位置
account = first_table.iloc[0, 1] 
branch = first_table.iloc[1, 1]

# 提取最后一张表格的日期、描述、金额字段
last_table = tables[-1].df
# 给最后一张表设置列名,根据实际表格字段调整
last_table.columns = ['序号', '日期', '描述', '借方金额', '贷方金额', '余额']
target_cols = last_table[['日期', '描述', '借方金额']] # 按需调整金额对应的列名

# 合并数据并存储为csv
result = pd.DataFrame({
    '账号': [account] * len(target_cols),
    '支行名称': [branch] * len(target_cols)
})
result = pd.concat([result, target_cols], axis=1)
result.to_csv(r"E:\user\programs\28_oct_bank_statement\提取结果.csv", index=False, encoding='utf_8_sig')

特殊情况适配(扫描版PDF)

如果完成Ghostscript安装后,camelot仍无法识别表格,说明目标PDF是扫描件,需使用OCR方案,先安装Tesseract OCR引擎并添加到系统环境变量,再安装依赖:
pip install pdf2image pytesseract pandas
对应的OCR提取代码可根据表格结构自定义规则提取目标字段。

内容的提问来源于stack exchange,提问作者Sand

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 01:15:05