You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:基于Python提取PDF表格以适配Google Vertex AI的LLM输入

解决方案

方法1:使用pdfplumber提取带位置信息的表格

pdfplumber可获取每个字符的坐标、对齐方式等元数据,适合处理表头对齐不一致的场景。你可以通过字符的y坐标判断是否属于同一行表头,手动修正结构:

import pdfplumber

def extract_table_with_pdfplumber(pdf_path):
    with pdfplumber.open(pdf_path) as pdf:
        first_page = pdf.pages[0]
        # 启用布局分析提取表格,调整容错度适配对齐差异
        table = first_page.extract_table(
            table_settings={
                "vertical_strategy": "lines",
                "horizontal_strategy": "lines",
                "snap_tolerance": 5,
            }
        )
        # 基于y坐标合并同一行的分散表头
        headers = []
        prev_y = None
        current_header = ""
        # 限定表头区域为页面顶部20%,可根据实际调整
        for cell in first_page.extract_words(x_tolerance=5, y_tolerance=5):
            if cell["top"] < first_page.height * 0.2:
                if prev_y is None or abs(cell["top"] - prev_y) < 3:
                    current_header += " " + cell["text"]
                else:
                    headers.append(current_header.strip())
                    current_header = cell["text"]
                prev_y = cell["top"]
        headers.append(current_header.strip())
        # 替换原表格第一行为修正后的表头
        if table:
            table[0] = headers
        return table

# 调用示例
result = extract_table_with_pdfplumber("your_doc_converted.pdf")
for row in result:
    print(row)

方法2:使用Camelot的Stream模式处理非严格网格表格

Camelot的Stream模式专门适配无明确网格线但有文本布局的表格,能更好识别对齐不同的表头:

import camelot

def extract_table_with_camelot(pdf_path):
    # 用Stream模式提取,增大行容错适配表头居中
    tables = camelot.read_pdf(
        pdf_path,
        flavor="stream",
        table_areas=["0,800,600,0"],  # 格式:x1,y1,x2,y2,需根据实际页面调整区域
        row_tol=10,
        column_tol=5
    )
    # 转换为结构化数据并修正表头
    df = tables[0].df
    df.columns = [col.strip() for col in df.iloc[0]]
    df = df[1:]
    return df.to_dict("records")

# 调用示例
result = extract_table_with_camelot("your_doc_converted.pdf")
for row in result:
    print(row)

方法3:优化Google Document AI的表格提取

若坚持使用Document AI,可改用Table Parser Processor而非通用OCR,同时配置表格提取参数优化结构化输出:

from google.cloud import documentai_v1 as documentai

def extract_table_with_document_ai(project_id, location, processor_id, pdf_path):
    client = documentai.DocumentProcessorServiceClient()
    name = client.processor_path(project_id, location, processor_id)

    with open(pdf_path, "rb") as image:
        image_content = image.read()

    raw_document = documentai.RawDocument(
        content=image_content, mime_type="application/pdf"
    )

    # 配置表格提取边界和参数
    table_extraction_config = documentai.TableExtractionConfig(
        enabled=True,
        table_bound_hints=[
            documentai.TableBoundHint(
                bounding_box=documentai.BoundingPoly(
                    vertices=[
                        documentai.Vertex(x=0, y=0),
                        documentai.Vertex(x=600, y=0),
                        documentai.Vertex(x=600, y=800),
                        documentai.Vertex(x=0, y=800),
                    ]
                )
            )
        ]
    )

    request = documentai.ProcessRequest(
        name=name,
        raw_document=raw_document,
        table_extraction_config=table_extraction_config
    )

    result = client.process_document(request=request)
    document = result.document

    # 解析表格结构化数据
    tables = []
    for page in document.pages:
        for table in page.tables:
            table_data = []
            # 提取表头行
            header_row = []
            for cell in table.header_rows[0].cells:
                text = "".join([seg.content for seg in cell.layout.text_anchor.text_segments])
                header_row.append(text.strip())
            table_data.append(header_row)
            # 提取内容行
            for row in table.body_rows:
                row_data = []
                for cell in row.cells:
                    text = "".join([seg.content for seg in cell.layout.text_anchor.text_segments])
                    row_data.append(text.strip())
                table_data.append(row_data)
            tables.append(table_data)
    return tables

# 调用示例(需提前配置GCP认证)
result = extract_table_with_document_ai(
    project_id="your-gcp-project",
    location="us",
    processor_id="your-table-processor-id",
    pdf_path="your_doc_converted.pdf"
)
for table in result:
    for row in table:
        print(row)

关键注意事项

  • 所有方法针对Doc转换的原生PDF,无需处理扫描件OCR问题,核心依赖文本位置信息而非像素分析。
  • 处理表头对齐问题的核心逻辑是通过字符/单元格的y坐标阈值判断是否属于同一行,可根据实际表格调整容错参数。
  • 可结合多个工具的结果交叉验证,比如用pdfplumber的位置信息修正Camelot的表头错误。

内容的提问来源于stack exchange,提问作者Sarthak Pan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 04:30:26