You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Camelot-py基于指定列文本拆分PDF表格行?

使用Camelot-py提取PDF表格的问题

我正尝试用Camelot-py库提取PDF里的表格信息,先后试了两种解析模式,都没得到预期结果:

Stream模式

一开始用stream模式,代码如下:

import camelot
tables = camelot.read_pdf('sample.pdf', flavor='stream', pages='1', columns=['According to the110,400'], split_text=True, row_tol=10)
tables.export('ipc_export.csv', f='csv', compress=True)
tables[0]
tables[0].parsing_report
tables[0].to_csv('ipc_export.csv')
tables[0].df

但不管怎么调整columns参数,都达不到预期效果。

Lattice模式

之后换成lattice模式,这个模式能准确识别列,但因为PDF源文件没有行分隔线,整个表格的内容都被提取到了同一行里。代码如下:

import camelot
tables = camelot.read_pdf('sample_camelot_extract.pdf', flavor='lattice', pages='1')
tables.export('ipc_export.csv', f='csv', compress=True)
tables[0]
tables[0].parsing_report
tables[0].to_csv('ipc_export.csv')
tables[0].df

我的核心需求是:当第一列(FIG ITEM)出现新文本时,就作为新行的起始。现在试过两种模式,但不确定哪种方案更适合解决这个问题。


内容的提问来源于stack exchange,提问作者KAmri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 20:40:30