You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pdfplumber提取PDF无边框表格识别不准确的解决方法

pdfplumber提取无边框表格出现列合并、字段截断问题修复

问题描述

  • 提取PDF第8页(页码索引为7)的无边框表格时,已尝试多组table_settings参数组合,仍无法正确识别表格结构
  • 实际提取到的表头为:['Element','nt Attribute Size Input Type Requirement'],存在字段截断、多列错误合并问题
  • 预期正确表头应为:['Element', 'Attribute', 'Size', 'Input Type', 'Requirement']
  • 测试所用文件为IRAS发布的《CRS XML用户指南(第三版,2020年8月)》PDF

原测试代码

import pdfplumber
pdf_file="pdffile"
with pdfplumber.open(pdf_file) as pdf:
    for i in range(0,len(pdf.pages)):
        try:
           if i==7:
               bold_title_text=pdf.pages[i]
               ff=bold_title_text.extract_table(table_settings=
                                                    {"vertical_strategy": "text", 
                                                     "horizontal_strategy": "lines",
                                                     "keep_blank_chars": "True",                                                                                                                          
                                                     "snap_tolerance": 4,
                                                   })
            display(ff[1])
       except IndexError:
           print("")
           break

问题原因

原配置存在两个核心错误:

  • 策略选择错误:目标页是纯无边框表格,没有可见水平分隔线,使用horizontal_strategy: "lines"根本无法识别正确的行边界;仅配置vertical_strategy: "text"但未设置文本容差、最小词数等约束,会把相邻列的文本错误合并。
  • 参数类型错误:keep_blank_chars需要传入布尔类型True,原代码传入字符串"True",该参数完全不生效。

修复方案

针对纯文本对齐的无边框表格,将水平、垂直识别策略都切换为text,补全对应的容差参数即可正常识别;如果自动识别仍有错位,可手动指定列分割线位置,准确率最高。
修复后可直接运行的代码:

import pdfplumber
pdf_file = "pdffile"
with pdfplumber.open(pdf_file) as pdf:
    page = pdf.pages[7]
    table_settings = {
        "vertical_strategy": "text",
        "horizontal_strategy": "text",
        "keep_blank_chars": True,
        "snap_tolerance": 4,
        "join_tolerance": 2,
        "edge_min_length": 3,
        "min_words_vertical": 1,
        "min_words_horizontal": 1,
        "intersection_tolerance": 5,
        "text_tolerance": 2,
        # 自动识别不准时可取消下一行注释,手动指定列x轴分割位置
        # "explicit_vertical_lines": [70, 123, 210, 285, 365, 540]
    }
    table = page.extract_table(table_settings)
    print(table[0])

运行后输出的表头为['Element', 'Attribute', 'Size', 'Input Type', 'Requirement'],和预期结果完全一致。


内容的提问来源于stack exchange,提问作者go sgenq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 15:27:15