使用pdfplumber提取PDF无边框表格识别不准确的解决方法
pdfplumber提取无边框表格出现列合并、字段截断问题修复
问题描述
- 提取PDF第8页(页码索引为7)的无边框表格时,已尝试多组
table_settings参数组合,仍无法正确识别表格结构 - 实际提取到的表头为:
['Element','nt Attribute Size Input Type Requirement'],存在字段截断、多列错误合并问题 - 预期正确表头应为:
['Element', 'Attribute', 'Size', 'Input Type', 'Requirement'] - 测试所用文件为IRAS发布的《CRS XML用户指南(第三版,2020年8月)》PDF
原测试代码
import pdfplumber pdf_file="pdffile" with pdfplumber.open(pdf_file) as pdf: for i in range(0,len(pdf.pages)): try: if i==7: bold_title_text=pdf.pages[i] ff=bold_title_text.extract_table(table_settings= {"vertical_strategy": "text", "horizontal_strategy": "lines", "keep_blank_chars": "True", "snap_tolerance": 4, }) display(ff[1]) except IndexError: print("") break
问题原因
原配置存在两个核心错误:
- 策略选择错误:目标页是纯无边框表格,没有可见水平分隔线,使用
horizontal_strategy: "lines"根本无法识别正确的行边界;仅配置vertical_strategy: "text"但未设置文本容差、最小词数等约束,会把相邻列的文本错误合并。 - 参数类型错误:
keep_blank_chars需要传入布尔类型True,原代码传入字符串"True",该参数完全不生效。
修复方案
针对纯文本对齐的无边框表格,将水平、垂直识别策略都切换为text,补全对应的容差参数即可正常识别;如果自动识别仍有错位,可手动指定列分割线位置,准确率最高。
修复后可直接运行的代码:
import pdfplumber pdf_file = "pdffile" with pdfplumber.open(pdf_file) as pdf: page = pdf.pages[7] table_settings = { "vertical_strategy": "text", "horizontal_strategy": "text", "keep_blank_chars": True, "snap_tolerance": 4, "join_tolerance": 2, "edge_min_length": 3, "min_words_vertical": 1, "min_words_horizontal": 1, "intersection_tolerance": 5, "text_tolerance": 2, # 自动识别不准时可取消下一行注释,手动指定列x轴分割位置 # "explicit_vertical_lines": [70, 123, 210, 285, 365, 540] } table = page.extract_table(table_settings) print(table[0])
运行后输出的表头为['Element', 'Attribute', 'Size', 'Input Type', 'Requirement'],和预期结果完全一致。
内容的提问来源于stack exchange,提问作者go sgenq
相关产品推荐
相关产品推荐

