如何用pdfplumber高效提取PDF稀疏表格?无法获取特定值求解
优化pdfplumber提取PDF表格的策略
核心方法:使用内置表格提取函数
pdfplumber自带的extract_table()/extract_tables()是专门针对表格结构设计的,比纯文本提取更精准,能直接返回二维列表格式的表格数据。
- 基础提取示例:
import pdfplumber with pdfplumber.open(doc) as pdf: page = pdf.pages[0] # 提取当前页所有表格,返回二维列表 tables = page.extract_tables() if tables: # 假设目标表格为页面第一个表格 target_table = tables[0] # 提取第一、二行并转为整数数组(过滤空单元格) size_array = [int(cell.strip()) for cell in target_table[0] if cell and cell.strip()] qty_array = [int(cell.strip()) for cell in target_table[1] if cell and cell.strip()] print(f"尺码数组: {size_array}") print(f"数量数组: {qty_array}")
优化表格识别精度
如果默认参数识别表格有偏差,通过table_settings调整识别逻辑,适配不同排版的表格:
- 参数调整示例:
table_settings = { "vertical_strategy": "text", # 基于文本对齐识别垂直列 "horizontal_strategy": "text", # 基于文本对齐识别水平行 "snap_tolerance": 5, # 文本与线条的对齐容忍度 "join_tolerance": 3 # 合并相邻线条的容忍度 } with pdfplumber.open(doc) as pdf: page = pdf.pages[0] tables = page.extract_tables(table_settings=table_settings) # 后续数据处理同上
合并非表格字段提取
你需要的结果包含表格外的字段(如编号、名称等),可以结合文本提取与正则匹配获取这些字段,再与表格数组合并:
import re with pdfplumber.open(doc) as pdf: page = pdf.pages[0] # 提取表格数据 tables = page.extract_tables() size_array = [int(cell.strip()) for cell in tables[0][0] if cell and cell.strip()] qty_array = [int(cell.strip()) for cell in tables[0][1] if cell and cell.strip()] # 提取整页文本并匹配非表格字段(需根据实际文本格式调整正则) full_text = page.extract_text(layout=True) # 假设字段格式为:编号 编码 名称 材质 颜色 总数量 单价 折扣 总价 pattern = r"(\d+)\s+(\w+)\s+(.*?)\s+(.*?)\s+(.*?)\s+(\d+)\s+(\d+\.\d+)\s+(\d+)\s+(\d+\.\d+)" match = re.search(pattern, full_text) if match: # 拆分匹配到的字段 base_fields = list(match.groups()) # 合并最终结果 final_result = base_fields[:5] + [size_array, qty_array] + base_fields[5:] print(",".join(map(str, final_result)))
可视化调试
用to_image()绘制表格识别区域,直观确认是否正确识别目标表格:
with pdfplumber.open(doc) as pdf: page = pdf.pages[0] im = page.to_image(resolution=400) # 绘制识别到的表格边框 im.draw_rects(page.find_tables()) im.show()
内容的提问来源于stack exchange,提问作者Herojos
相关产品推荐
相关产品推荐

