You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用pdfplumber高效提取PDF稀疏表格?无法获取特定值求解

优化pdfplumber提取PDF表格的策略

核心方法:使用内置表格提取函数

pdfplumber自带的extract_table()/extract_tables()是专门针对表格结构设计的,比纯文本提取更精准,能直接返回二维列表格式的表格数据。

  • 基础提取示例:
import pdfplumber

with pdfplumber.open(doc) as pdf:
    page = pdf.pages[0]
    # 提取当前页所有表格,返回二维列表
    tables = page.extract_tables()
    if tables:
        # 假设目标表格为页面第一个表格
        target_table = tables[0]
        # 提取第一、二行并转为整数数组(过滤空单元格)
        size_array = [int(cell.strip()) for cell in target_table[0] if cell and cell.strip()]
        qty_array = [int(cell.strip()) for cell in target_table[1] if cell and cell.strip()]
        print(f"尺码数组: {size_array}")
        print(f"数量数组: {qty_array}")

优化表格识别精度

如果默认参数识别表格有偏差,通过table_settings调整识别逻辑,适配不同排版的表格:

  • 参数调整示例:
table_settings = {
    "vertical_strategy": "text",  # 基于文本对齐识别垂直列
    "horizontal_strategy": "text", # 基于文本对齐识别水平行
    "snap_tolerance": 5,  # 文本与线条的对齐容忍度
    "join_tolerance": 3   # 合并相邻线条的容忍度
}

with pdfplumber.open(doc) as pdf:
    page = pdf.pages[0]
    tables = page.extract_tables(table_settings=table_settings)
    # 后续数据处理同上

合并非表格字段提取

你需要的结果包含表格外的字段(如编号、名称等),可以结合文本提取与正则匹配获取这些字段,再与表格数组合并:

import re

with pdfplumber.open(doc) as pdf:
    page = pdf.pages[0]
    # 提取表格数据
    tables = page.extract_tables()
    size_array = [int(cell.strip()) for cell in tables[0][0] if cell and cell.strip()]
    qty_array = [int(cell.strip()) for cell in tables[0][1] if cell and cell.strip()]
    # 提取整页文本并匹配非表格字段(需根据实际文本格式调整正则)
    full_text = page.extract_text(layout=True)
    # 假设字段格式为:编号 编码 名称 材质 颜色 总数量 单价 折扣 总价
    pattern = r"(\d+)\s+(\w+)\s+(.*?)\s+(.*?)\s+(.*?)\s+(\d+)\s+(\d+\.\d+)\s+(\d+)\s+(\d+\.\d+)"
    match = re.search(pattern, full_text)
    
    if match:
        # 拆分匹配到的字段
        base_fields = list(match.groups())
        # 合并最终结果
        final_result = base_fields[:5] + [size_array, qty_array] + base_fields[5:]
        print(",".join(map(str, final_result)))

可视化调试

用to_image()绘制表格识别区域,直观确认是否正确识别目标表格:

with pdfplumber.open(doc) as pdf:
    page = pdf.pages[0]
    im = page.to_image(resolution=400)
    # 绘制识别到的表格边框
    im.draw_rects(page.find_tables())
    im.show()

内容的提问来源于stack exchange,提问作者Herojos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 13:07:42