如何用Python基于动态坐标从PDF提取文本至CSV
动态提取PDF文本(无需手动指定坐标)
原代码通过硬编码PDF坐标提取HBL文件中的文本并导出CSV,但不同PDF的布局差异会导致坐标失效,无法通用。下面是改进方案,直接从PDF生成的XML中通过文本标签动态定位内容,无需手动输入坐标:
改进后的代码
import pdfquery import pandas as pd def extract_field_by_label(pdf, label_text, direction="right", tolerance=50): """根据标签文本动态提取对应内容""" # 定位标签元素 label = pdf.pq(f'LTTextLineHorizontal:contains("{label_text}")') if not label: return "" # 获取标签的坐标 label_bbox = label.attr('bbox').split(',') label_x1 = float(label_bbox[2]) label_y0 = float(label_bbox[1]) label_y1 = float(label_bbox[3]) # 根据方向构建查询,匹配标签相邻区域的文本 if direction == "right": query = f'LTTextLineHorizontal:gt("{label_x1}", 0):between("{label_y0 - tolerance}", "{label_y1 + tolerance}")' elif direction == "below": query = f'LTTextLineHorizontal:lt("{label_y0}", 1):between("{float(label_bbox[0]) - tolerance}", "{float(label_bbox[2]) + tolerance}")' # 合并多行内容 elements = pdf.pq(query) return ' '.join([elem.text() for elem in elements]) def pdfscrape(pdf): # 通过标签动态提取各字段 shipper = extract_field_by_label(pdf, "Shipper", direction="below", tolerance=100) consignee = extract_field_by_label(pdf, "Consignee", direction="below", tolerance=100) notify_party = extract_field_by_label(pdf, "Notify Party", direction="below", tolerance=50) ocean_vessel = extract_field_by_label(pdf, "Ocean Vessel", direction="right") port_of_loading = extract_field_by_label(pdf, "Port of Loading", direction="right") port_of_discharge = extract_field_by_label(pdf, "Port of Discharge", direction="right") place_of_delivery = extract_field_by_label(pdf, "Place of Delivery", direction="right") for_delivery_of_goods = extract_field_by_label(pdf, "For Delivery of Goods", direction="below", tolerance=100) container_no_and_seal_no = extract_field_by_label(pdf, "Container No./Seal No.", direction="right") no_of_container = extract_field_by_label(pdf, "No. of Container", direction="right") gross_weight = extract_field_by_label(pdf, "Gross Weight", direction="right") measurement = extract_field_by_label(pdf, "Measurement", direction="right") # 组合成DataFrame return pd.DataFrame({ 'Shipper': [shipper], 'Consignee': [consignee], 'Notify_party': [notify_party], 'Ocean_vessel': [ocean_vessel], 'Port_of_loading': [port_of_loading], 'Port_of_discharge': [port_of_discharge], 'Place_of_delivery': [place_of_delivery], 'For_delivery_of_goods': [for_delivery_of_goods], 'Container_no_and_seal_no': [container_no_and_seal_no], 'No_of_container': [no_of_container], 'Gross_weight': [gross_weight], 'Measurement': [measurement] }) # 主流程 pdf = pdfquery.PDFQuery('HBL.PDF') pagecount = pdf.doc.catalog['Pages'].resolve()['Count'] master = pd.DataFrame() for p in range(pagecount): pdf.load(p) page_data = pdfscrape(pdf) master = pd.concat([master, page_data], ignore_index=True) # 导出CSV master.to_csv('output.csv', index=False) # 可选:保存XML用于调试 pdf.tree.write('pdfXML.xml', pretty_print=True)
核心改进说明
- 动态定位逻辑:通过
extract_field_by_label函数先找到字段对应的标签文本(如"Shipper"),再根据标签坐标自动定位相邻区域的内容,彻底摆脱硬编码坐标的限制。 - 适配不同布局:支持
right(标签右侧)和below(标签下方)两种方向,通过tolerance参数调整坐标匹配范围,适配不同PDF的布局差异。 - 多行内容合并:自动合并同一字段的多行文本,避免内容拆分导致的信息缺失。
- 通用扩展性:只需调整标签文本和方向参数,即可适配不同格式的HBL或其他结构化PDF文件。
内容的提问来源于stack exchange,提问作者Abhinav Bhardwaj
相关产品推荐
相关产品推荐

