You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Fitz的PDF表格文本提取:行列区分与可解读输出方案问询

优化Fitz提取PDF表格文本的结构化方案

针对你用Fitz提取PDF表格文本时,无法明确数值对应列的问题,以下是几种更通用的解决方案:

1. 按单元格级边界框精准提取

不要仅用整个表格的边界框,而是定义每个列标题和单元格的独立坐标,分别提取内容后直接关联对应关系。这种方式完全依赖坐标定位,不受字体、间距差异影响。

示例优化代码:

import fitz

def extract_structured_table(pdf_file, header_coords, row_coords_list):
    pdf_document = fitz.open(pdf_file)
    table_data = []
    
    # 提取列标题
    headers = []
    for coord in header_coords:
        x1, y1, x2, y2 = map(float, coord.split(','))
        rect = fitz.Rect(x1, y1, x2, y2)
        header_text = pdf_document[0].get_text("text", clip=rect).strip()
        headers.append(header_text)
    
    # 提取每一行的单元格内容
    for row_coords in row_coords_list:
        row_data = {}
        for idx, coord in enumerate(row_coords):
            x1, y1, x2, y2 = map(float, coord.split(','))
            rect = fitz.Rect(x1, y1, x2, y2)
            cell_text = pdf_document[0].get_text("text", clip=rect).strip()
            row_data[headers[idx]] = cell_text
        table_data.append(row_data)
    
    pdf_document.close()
    return table_data

# 使用示例:假设表头有3个单元格,两行数据各3个单元格
header_coords = ["x1,y1,x2,y2", "x3,y3,x4,y4", "x5,y5,x6,y6"]
row_coords_list = [
    ["x1,y7,x2,y8", "x3,y7,x4,y8", "x5,y7,x6,y8"],
    ["x1,y9,x2,y10", "x3,y9,x4,y10", "x5,y9,x6,y10"]
]
structured_table = extract_structured_table("your_pdf.pdf", header_coords, row_coords_list)
# 输出结构化结果
for row in structured_table:
    print(row)

2. 利用文本块的位置信息自动分组

通过Fitz的get_text("words")获取每个文本片段的坐标(x0, y0, x1, y1),然后根据坐标对文本进行行、列分组,无需预先定义所有单元格坐标:

  • 按y坐标分组:y值相近的文本属于同一行(可设置小阈值,比如5)
  • 按x坐标排序:同一行内x值从小到大对应从左到右的列
  • 结合表头行的位置,将后续行的文本与表头关联

示例代码:

import fitz
from collections import defaultdict

def extract_table_via_positions(pdf_file, table_rect):
    pdf_document = fitz.open(pdf_file)
    page = pdf_document[0]
    x1, y1, x2, y2 = map(float, table_rect.split(','))
    table_area = fitz.Rect(x1, y1, x2, y2)
    
    # 获取表格区域内的所有文本块,每个元素格式:(x0, y0, x1, y1, text, block_no, line_no, word_no)
    words = page.get_text("words", clip=table_area)
    
    # 按行分组:根据y坐标聚类
    row_threshold = 5
    rows = defaultdict(list)
    for word in words:
        y0 = word[1]
        # 匹配同一行的y坐标组
        matched_row = None
        for row_y in rows.keys():
            if abs(y0 - row_y) < row_threshold:
                matched_row = row_y
                break
        if matched_row:
            rows[matched_row].append(word)
        else:
            rows[y0].append(word)
    
    # 按y坐标排序行,从上到下
    sorted_rows = sorted(rows.items(), key=lambda item: item[0])
    
    # 每行内按x坐标排序文本,从左到右
    table_content = []
    for _, row_words in sorted_rows:
        sorted_row = sorted(row_words, key=lambda w: w[0])
        row_text = [w[4].strip() for w in sorted_row]
        table_content.append(row_text)
    
    # 转换为结构化字典
    headers = table_content[0]
    data_rows = table_content[1:]
    structured_data = [dict(zip(headers, row)) for row in data_rows]
    
    pdf_document.close()
    return structured_data

# 使用示例:传入表格的整体边界框坐标
structured_data = extract_table_via_positions("your_pdf.pdf", "x1,y1,x2,y2")
# 输出Markdown格式表格增强可读性
print("| " + " | ".join(structured_data[0].keys()) + " |")
print("| " + " | ".join(["---"]*len(structured_data[0].keys())) + " |")
for row in structured_data:
    print("| " + " | ".join(row.values()) + " |")

3. 直接输出可读格式

提取完成后,将结果转换为Markdown表格或CSV格式,让阅读者能直观看到列与值的对应关系。比如用Python的csv模块实现CSV输出:

import csv

def save_to_csv(data, filename):
    with open(filename, 'w', newline='', encoding='utf-8') as f:
        writer = csv.DictWriter(f, fieldnames=data[0].keys())
        writer.writeheader()
        writer.writerows(data)

# 调用示例
save_to_csv(structured_data, "table_output.csv")

内容的提问来源于stack exchange,提问作者Apoorva

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 04:02:47