You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pdfplumber提取含空单元格无框线PDF表格遇提取异常求助

解决方案

方法1:优化pdfplumber表格识别参数

默认extract_tables()可能因合并表头、表格边界模糊导致识别不全,可通过调整table_settings参数强制识别完整表格:

import pdfplumber
import pandas as pd

pdf_file = 'D:/Input/Book1.pdf'
with pdfplumber.open(pdf_file) as pdf:
    page = pdf.pages[0]
    # 自定义表格识别规则,优先按线条识别行列,也可手动指定表格边界框
    table_settings = {
        "vertical_strategy": "lines",
        "horizontal_strategy": "lines",
        # 若已知表格坐标范围,可手动设置bbox,比如通过page.debug_tablefinder()查看坐标
        # "bbox": (50, 100, 750, 500)
    }
    tables = page.extract_tables(table_settings=table_settings)
    
    # 合并两行表头为完整列名
    header_row1 = tables[0][0]
    header_row2 = tables[0][1]
    headers = []
    for h1, h2 in zip(header_row1, header_row2):
        headers.append(f"{h1} {h2}".strip() if h2 else h1.strip())
    
    # 提取数据行并生成DataFrame
    data_rows = tables[0][2:]
    df = pd.DataFrame(data_rows, columns=headers)
    print(df)

方法2:基于文本位置的精准分割(适配空单元格丢失问题)

如果表格识别始终失效,可通过extract_words()获取文本块坐标,按列的水平位置分组补全空单元格:

import pdfplumber
import pandas as pd
from collections import defaultdict

pdf_file = 'D:/Input/Book1.pdf'
with pdfplumber.open(pdf_file) as pdf:
    page = pdf.pages[0]
    # 提取带坐标的文本块,容错相邻文本块的位置误差
    words = page.extract_words(x_tolerance=5, y_tolerance=5)
    
    # 按y坐标分组,将同一行的文本块归为一组
    rows = defaultdict(list)
    for word in words:
        row_key = round(word['top'], 0)  # 用y坐标整数部分避免浮点误差
        rows[row_key].append(word)
    
    # 每行文本块按x坐标从左到右排序
    sorted_rows = [sorted(rows[row], key=lambda x: x['x0']) for row in sorted(rows.keys())]
    
    # 手动指定列的水平边界(需根据实际PDF列位置调整)
    col_bounds = [0, 80, 150, 220, 300, 380, 460, 540, 620, 700]
    
    # 按列边界分割每行数据,空单元格补pd.NA
    data = []
    for row in sorted_rows:
        row_data = []
        for left, right in zip(col_bounds[:-1], col_bounds[1:]):
            cell_text = [w['text'] for w in row if left <= w['x0'] < right]
            row_data.append(' '.join(cell_text) if cell_text else pd.NA)
        data.append(row_data)
    
    # 合并表头并生成DataFrame
    header_row1 = data[1]
    header_row2 = data[2]
    headers = [f"{h1} {h2}".strip() if h2 else h1.strip() for h1, h2 in zip(header_row1, header_row2)]
    data_rows = data[3:]
    df = pd.DataFrame(data_rows, columns=headers)
    print(df)

关键提示

  • 方法1可先运行page.debug_tablefinder()查看pdfplumber识别的表格区域,针对性调整参数;
  • 方法2的col_bounds可通过page.to_image().draw_rects(page.extract_words())可视化文本块位置后确定。

内容的提问来源于stack exchange,提问作者Roshan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 01:40:15