求助:基于Python提取PDF表格以适配Google Vertex AI的LLM输入
解决方案
方法1:使用pdfplumber提取带位置信息的表格
pdfplumber可获取每个字符的坐标、对齐方式等元数据,适合处理表头对齐不一致的场景。你可以通过字符的y坐标判断是否属于同一行表头,手动修正结构:
import pdfplumber def extract_table_with_pdfplumber(pdf_path): with pdfplumber.open(pdf_path) as pdf: first_page = pdf.pages[0] # 启用布局分析提取表格,调整容错度适配对齐差异 table = first_page.extract_table( table_settings={ "vertical_strategy": "lines", "horizontal_strategy": "lines", "snap_tolerance": 5, } ) # 基于y坐标合并同一行的分散表头 headers = [] prev_y = None current_header = "" # 限定表头区域为页面顶部20%,可根据实际调整 for cell in first_page.extract_words(x_tolerance=5, y_tolerance=5): if cell["top"] < first_page.height * 0.2: if prev_y is None or abs(cell["top"] - prev_y) < 3: current_header += " " + cell["text"] else: headers.append(current_header.strip()) current_header = cell["text"] prev_y = cell["top"] headers.append(current_header.strip()) # 替换原表格第一行为修正后的表头 if table: table[0] = headers return table # 调用示例 result = extract_table_with_pdfplumber("your_doc_converted.pdf") for row in result: print(row)
方法2:使用Camelot的Stream模式处理非严格网格表格
Camelot的Stream模式专门适配无明确网格线但有文本布局的表格,能更好识别对齐不同的表头:
import camelot def extract_table_with_camelot(pdf_path): # 用Stream模式提取,增大行容错适配表头居中 tables = camelot.read_pdf( pdf_path, flavor="stream", table_areas=["0,800,600,0"], # 格式:x1,y1,x2,y2,需根据实际页面调整区域 row_tol=10, column_tol=5 ) # 转换为结构化数据并修正表头 df = tables[0].df df.columns = [col.strip() for col in df.iloc[0]] df = df[1:] return df.to_dict("records") # 调用示例 result = extract_table_with_camelot("your_doc_converted.pdf") for row in result: print(row)
方法3:优化Google Document AI的表格提取
若坚持使用Document AI,可改用Table Parser Processor而非通用OCR,同时配置表格提取参数优化结构化输出:
from google.cloud import documentai_v1 as documentai def extract_table_with_document_ai(project_id, location, processor_id, pdf_path): client = documentai.DocumentProcessorServiceClient() name = client.processor_path(project_id, location, processor_id) with open(pdf_path, "rb") as image: image_content = image.read() raw_document = documentai.RawDocument( content=image_content, mime_type="application/pdf" ) # 配置表格提取边界和参数 table_extraction_config = documentai.TableExtractionConfig( enabled=True, table_bound_hints=[ documentai.TableBoundHint( bounding_box=documentai.BoundingPoly( vertices=[ documentai.Vertex(x=0, y=0), documentai.Vertex(x=600, y=0), documentai.Vertex(x=600, y=800), documentai.Vertex(x=0, y=800), ] ) ) ] ) request = documentai.ProcessRequest( name=name, raw_document=raw_document, table_extraction_config=table_extraction_config ) result = client.process_document(request=request) document = result.document # 解析表格结构化数据 tables = [] for page in document.pages: for table in page.tables: table_data = [] # 提取表头行 header_row = [] for cell in table.header_rows[0].cells: text = "".join([seg.content for seg in cell.layout.text_anchor.text_segments]) header_row.append(text.strip()) table_data.append(header_row) # 提取内容行 for row in table.body_rows: row_data = [] for cell in row.cells: text = "".join([seg.content for seg in cell.layout.text_anchor.text_segments]) row_data.append(text.strip()) table_data.append(row_data) tables.append(table_data) return tables # 调用示例(需提前配置GCP认证) result = extract_table_with_document_ai( project_id="your-gcp-project", location="us", processor_id="your-table-processor-id", pdf_path="your_doc_converted.pdf" ) for table in result: for row in table: print(row)
关键注意事项
- 所有方法针对Doc转换的原生PDF,无需处理扫描件OCR问题,核心依赖文本位置信息而非像素分析。
- 处理表头对齐问题的核心逻辑是通过字符/单元格的y坐标阈值判断是否属于同一行,可根据实际表格调整容错参数。
- 可结合多个工具的结果交叉验证,比如用pdfplumber的位置信息修正Camelot的表头错误。
内容的提问来源于stack exchange,提问作者Sarthak Pan
相关产品推荐
相关产品推荐

