基于Fitz的PDF表格文本提取:行列区分与可解读输出方案问询
优化Fitz提取PDF表格文本的结构化方案
针对你用Fitz提取PDF表格文本时,无法明确数值对应列的问题,以下是几种更通用的解决方案:
1. 按单元格级边界框精准提取
不要仅用整个表格的边界框,而是定义每个列标题和单元格的独立坐标,分别提取内容后直接关联对应关系。这种方式完全依赖坐标定位,不受字体、间距差异影响。
示例优化代码:
import fitz def extract_structured_table(pdf_file, header_coords, row_coords_list): pdf_document = fitz.open(pdf_file) table_data = [] # 提取列标题 headers = [] for coord in header_coords: x1, y1, x2, y2 = map(float, coord.split(',')) rect = fitz.Rect(x1, y1, x2, y2) header_text = pdf_document[0].get_text("text", clip=rect).strip() headers.append(header_text) # 提取每一行的单元格内容 for row_coords in row_coords_list: row_data = {} for idx, coord in enumerate(row_coords): x1, y1, x2, y2 = map(float, coord.split(',')) rect = fitz.Rect(x1, y1, x2, y2) cell_text = pdf_document[0].get_text("text", clip=rect).strip() row_data[headers[idx]] = cell_text table_data.append(row_data) pdf_document.close() return table_data # 使用示例:假设表头有3个单元格,两行数据各3个单元格 header_coords = ["x1,y1,x2,y2", "x3,y3,x4,y4", "x5,y5,x6,y6"] row_coords_list = [ ["x1,y7,x2,y8", "x3,y7,x4,y8", "x5,y7,x6,y8"], ["x1,y9,x2,y10", "x3,y9,x4,y10", "x5,y9,x6,y10"] ] structured_table = extract_structured_table("your_pdf.pdf", header_coords, row_coords_list) # 输出结构化结果 for row in structured_table: print(row)
2. 利用文本块的位置信息自动分组
通过Fitz的get_text("words")获取每个文本片段的坐标(x0, y0, x1, y1),然后根据坐标对文本进行行、列分组,无需预先定义所有单元格坐标:
- 按y坐标分组:y值相近的文本属于同一行(可设置小阈值,比如5)
- 按x坐标排序:同一行内x值从小到大对应从左到右的列
- 结合表头行的位置,将后续行的文本与表头关联
示例代码:
import fitz from collections import defaultdict def extract_table_via_positions(pdf_file, table_rect): pdf_document = fitz.open(pdf_file) page = pdf_document[0] x1, y1, x2, y2 = map(float, table_rect.split(',')) table_area = fitz.Rect(x1, y1, x2, y2) # 获取表格区域内的所有文本块,每个元素格式:(x0, y0, x1, y1, text, block_no, line_no, word_no) words = page.get_text("words", clip=table_area) # 按行分组:根据y坐标聚类 row_threshold = 5 rows = defaultdict(list) for word in words: y0 = word[1] # 匹配同一行的y坐标组 matched_row = None for row_y in rows.keys(): if abs(y0 - row_y) < row_threshold: matched_row = row_y break if matched_row: rows[matched_row].append(word) else: rows[y0].append(word) # 按y坐标排序行,从上到下 sorted_rows = sorted(rows.items(), key=lambda item: item[0]) # 每行内按x坐标排序文本,从左到右 table_content = [] for _, row_words in sorted_rows: sorted_row = sorted(row_words, key=lambda w: w[0]) row_text = [w[4].strip() for w in sorted_row] table_content.append(row_text) # 转换为结构化字典 headers = table_content[0] data_rows = table_content[1:] structured_data = [dict(zip(headers, row)) for row in data_rows] pdf_document.close() return structured_data # 使用示例:传入表格的整体边界框坐标 structured_data = extract_table_via_positions("your_pdf.pdf", "x1,y1,x2,y2") # 输出Markdown格式表格增强可读性 print("| " + " | ".join(structured_data[0].keys()) + " |") print("| " + " | ".join(["---"]*len(structured_data[0].keys())) + " |") for row in structured_data: print("| " + " | ".join(row.values()) + " |")
3. 直接输出可读格式
提取完成后,将结果转换为Markdown表格或CSV格式,让阅读者能直观看到列与值的对应关系。比如用Python的csv模块实现CSV输出:
import csv def save_to_csv(data, filename): with open(filename, 'w', newline='', encoding='utf-8') as f: writer = csv.DictWriter(f, fieldnames=data[0].keys()) writer.writeheader() writer.writerows(data) # 调用示例 save_to_csv(structured_data, "table_output.csv")
内容的提问来源于stack exchange,提问作者Apoorva
相关产品推荐
相关产品推荐

