pdfplumber提取无边框PDF表格失败问题求助
问题:无表格边框的PDF无法用pdfplumber提取表格数据,调用参数报错
我用reportlab生成了两个单页PDF文件,各包含一个表格,表格源数据如下:
data1 = [['00', '', '02', '', '04'], ['', '11', '', '13', ''], ['20', '', '22', '23', '24'], ['30', '31', '32', '', '34']]
需求是提取包含空单元格的完整表格行。带边框的PDF(pdf2)用下述代码可正常提取,但无边框的PDF(pdf1)无法获取表格数据:
with pdfplumber.open(path2pdf + savename1) as pdf1: # 获取第一页 page = pdf1.pages[0] # 提取页面文本 text = page.extract_text() # 提取所有表格数据 tables = page.extract_tables() # 遍历表格 for t_index in range(len(tables)): table = tables[t_index] # 遍历每行数据 for data in table: print(data)
将代码中的pdf1替换为pdf2即可得到预期结果。
后续尝试调整提取策略,使用以下代码时出现报错:
pdf_table = page.extract_tables(vertical_strategy='text', horizontal_strategy='text')
报错信息:
Traceback (most recent call last): File "/usr/lib/python3.8/idlelib/run.py", line 559, in runcode exec(code, self.locals) File "<pyshell#70>", line 1, in <module> TypeError: extract_tables() got an unexpected keyword argument 'vertical_strategy'
请问这是什么原因,该如何解决?
原因分析与解决方法
1. 参数报错原因
extract_tables()(复数形式,提取多表格)不支持直接传入vertical_strategy和horizontal_strategy参数——这两个参数属于extract_table()(单数形式,提取单个表格)的配置项,或是需要通过set_table_settings()方法提前设置表格识别规则,不能直接传给extract_tables()。
2. 无边框PDF表格提取解决方案
pdfplumber默认通过表格边框识别行列,无边框表格需要切换为基于文本对齐的识别策略,有两种可行方法:
方法一:针对单表格场景使用extract_table()
如果页面只有一个表格,直接调用extract_table()并指定策略参数即可:
with pdfplumber.open(path2pdf + savename1) as pdf1: page = pdf1.pages[0] # 基于文本对齐识别行列,适配无边框表格 table = page.extract_table(vertical_strategy="text", horizontal_strategy="text") for row in table: print(row)
方法二:针对多表格场景先设置识别规则
如果页面可能存在多个表格,先通过set_table_settings()配置识别策略,再调用extract_tables():
with pdfplumber.open(path2pdf + savename1) as pdf1: page = pdf1.pages[0] # 全局设置表格识别策略 page.set_table_settings({ "vertical_strategy": "text", "horizontal_strategy": "text" }) tables = page.extract_tables() for table in tables: for row in table: print(row)
3. 策略参数说明
vertical_strategy和horizontal_strategy的可选值:
lines:依赖PDF中的线条(默认值,适合带边框表格)text:依赖文本块的行列对齐关系识别表格(适合无边框表格)explicit:自定义行列坐标进行识别
内容的提问来源于stack exchange,提问作者Pedroski
相关产品推荐
相关产品推荐

