You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的Camelot捕获PDF表格完整尺寸并转换?

问题解决:Camelot捕获PDF完整表格的调整方案

你当前代码的核心问题是没有正确设置area参数,以及对Camelot的坐标规则不匹配,导致无法覆盖整页识别。以下是修正后的方案:

代码修正点

  1. 正确赋值area参数:你之前把page_area设为[0,0,0,0],相当于没有指定有效识别区域。Camelot的area参数格式为[top, left, bottom, right],坐标原点在PDF页面左下角,y轴向上,所以整页区域需要用获取到的页面宽高来赋值。
  2. 页码索引兼容:Camelot的pages参数默认从1开始计数,而PyPDF2的页面索引从0开始,需要对应调整避免读取错误页面。
  3. PyPDF2版本适配:新版本PyPDF2已弃用PdfFileReader,建议改用PdfReader,同时兼容旧版本写法。

修正后的完整代码

import camelot
import PyPDF2

# 读取PDF页面尺寸
pdf_file = open(r'C:\Users\PC\PycharmProjects\finstate.pdf', 'rb')
# 兼容新旧PyPDF2版本
try:
    pdf_reader = PyPDF2.PdfReader(pdf_file)
    page = pdf_reader.pages[10]  # 新版本用pages索引
except AttributeError:
    pdf_reader = PyPDF2.PdfFileReader(pdf_file)
    page = pdf_reader.getPage(10)

width = page.mediaBox.getWidth()
height = page.mediaBox.getHeight()
print("Width:", width)
print("Height:", height)

# 配置整页识别区域:[top, left, bottom, right]
# 页面顶部y坐标为height,底部为0;左x为0,右x为width
page_area = [height, 0, 0, width]

# 读取对应页面(PyPDF2的第10页对应Camelot的第11页)
pdf = camelot.read_pdf(
    r'C:\Users\PC\PycharmProjects\finstate.pdf',
    pages='11',
    flavor='stream',
    area=page_area
)

# 取该页的第一个表格(若页面仅一个表格)
first_table = pdf[0]
print(first_table.df)
first_table.to_csv(r'C:\Users\PC\Desktop\table.csv')

额外优化建议

  • 如果stream模式识别仍有列错位问题,可添加columns参数手动指定列的x坐标(如columns=[50, 150, 250, 350]),辅助Camelot拆分列。
  • 用camelot.plot(first_table, kind='grid').show()可视化识别结果,便于调整area或columns参数。

内容的提问来源于stack exchange,提问作者Jagwire

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 20:05:30