You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pdfplumber提取无边框PDF表格失败问题求助

问题:无表格边框的PDF无法用pdfplumber提取表格数据,调用参数报错

我用reportlab生成了两个单页PDF文件,各包含一个表格,表格源数据如下:

data1 = [['00', '', '02', '', '04'],  
['', '11', '', '13', ''],  
['20', '', '22', '23', '24'],  
['30', '31', '32', '', '34']]  

需求是提取包含空单元格的完整表格行。带边框的PDF(pdf2)用下述代码可正常提取,但无边框的PDF(pdf1)无法获取表格数据:

with pdfplumber.open(path2pdf + savename1) as pdf1:
    # 获取第一页
    page = pdf1.pages[0]
    # 提取页面文本
    text = page.extract_text()
    # 提取所有表格数据
    tables = page.extract_tables()
    # 遍历表格
    for t_index in range(len(tables)):
        table = tables[t_index]
        # 遍历每行数据
        for data in table:
            print(data)

将代码中的pdf1替换为pdf2即可得到预期结果。

后续尝试调整提取策略,使用以下代码时出现报错:

pdf_table = page.extract_tables(vertical_strategy='text', horizontal_strategy='text')

报错信息:

Traceback (most recent call last):
  File "/usr/lib/python3.8/idlelib/run.py", line 559, in runcode
    exec(code, self.locals)
  File "<pyshell#70>", line 1, in <module>
TypeError: extract_tables() got an unexpected keyword argument 'vertical_strategy'

请问这是什么原因,该如何解决?


原因分析与解决方法

1. 参数报错原因

extract_tables()(复数形式,提取多表格)不支持直接传入vertical_strategy和horizontal_strategy参数——这两个参数属于extract_table()(单数形式,提取单个表格)的配置项,或是需要通过set_table_settings()方法提前设置表格识别规则,不能直接传给extract_tables()。

2. 无边框PDF表格提取解决方案

pdfplumber默认通过表格边框识别行列,无边框表格需要切换为基于文本对齐的识别策略,有两种可行方法:

方法一:针对单表格场景使用extract_table()

如果页面只有一个表格,直接调用extract_table()并指定策略参数即可:

with pdfplumber.open(path2pdf + savename1) as pdf1:
    page = pdf1.pages[0]
    # 基于文本对齐识别行列,适配无边框表格
    table = page.extract_table(vertical_strategy="text", horizontal_strategy="text")
    for row in table:
        print(row)

方法二:针对多表格场景先设置识别规则

如果页面可能存在多个表格,先通过set_table_settings()配置识别策略,再调用extract_tables():

with pdfplumber.open(path2pdf + savename1) as pdf1:
    page = pdf1.pages[0]
    # 全局设置表格识别策略
    page.set_table_settings({
        "vertical_strategy": "text",
        "horizontal_strategy": "text"
    })
    tables = page.extract_tables()
    for table in tables:
        for row in table:
            print(row)

3. 策略参数说明

vertical_strategy和horizontal_strategy的可选值:

  • lines:依赖PDF中的线条(默认值,适合带边框表格)
  • text:依赖文本块的行列对齐关系识别表格(适合无边框表格)
  • explicit:自定义行列坐标进行识别

内容的提问来源于stack exchange,提问作者Pedroski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 07:15:59