You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python基于动态坐标从PDF提取文本至CSV

动态提取PDF文本(无需手动指定坐标)

原代码通过硬编码PDF坐标提取HBL文件中的文本并导出CSV,但不同PDF的布局差异会导致坐标失效,无法通用。下面是改进方案,直接从PDF生成的XML中通过文本标签动态定位内容,无需手动输入坐标:

改进后的代码

import pdfquery
import pandas as pd

def extract_field_by_label(pdf, label_text, direction="right", tolerance=50):
    """根据标签文本动态提取对应内容"""
    # 定位标签元素
    label = pdf.pq(f'LTTextLineHorizontal:contains("{label_text}")')
    if not label:
        return ""
    
    # 获取标签的坐标
    label_bbox = label.attr('bbox').split(',')
    label_x1 = float(label_bbox[2])
    label_y0 = float(label_bbox[1])
    label_y1 = float(label_bbox[3])
    
    # 根据方向构建查询,匹配标签相邻区域的文本
    if direction == "right":
        query = f'LTTextLineHorizontal:gt("{label_x1}", 0):between("{label_y0 - tolerance}", "{label_y1 + tolerance}")'
    elif direction == "below":
        query = f'LTTextLineHorizontal:lt("{label_y0}", 1):between("{float(label_bbox[0]) - tolerance}", "{float(label_bbox[2]) + tolerance}")'
    
    # 合并多行内容
    elements = pdf.pq(query)
    return ' '.join([elem.text() for elem in elements])

def pdfscrape(pdf):
    # 通过标签动态提取各字段
    shipper = extract_field_by_label(pdf, "Shipper", direction="below", tolerance=100)
    consignee = extract_field_by_label(pdf, "Consignee", direction="below", tolerance=100)
    notify_party = extract_field_by_label(pdf, "Notify Party", direction="below", tolerance=50)
    ocean_vessel = extract_field_by_label(pdf, "Ocean Vessel", direction="right")
    port_of_loading = extract_field_by_label(pdf, "Port of Loading", direction="right")
    port_of_discharge = extract_field_by_label(pdf, "Port of Discharge", direction="right")
    place_of_delivery = extract_field_by_label(pdf, "Place of Delivery", direction="right")
    for_delivery_of_goods = extract_field_by_label(pdf, "For Delivery of Goods", direction="below", tolerance=100)
    container_no_and_seal_no = extract_field_by_label(pdf, "Container No./Seal No.", direction="right")
    no_of_container = extract_field_by_label(pdf, "No. of Container", direction="right")
    gross_weight = extract_field_by_label(pdf, "Gross Weight", direction="right")
    measurement = extract_field_by_label(pdf, "Measurement", direction="right")
    
    # 组合成DataFrame
    return pd.DataFrame({
        'Shipper': [shipper],
        'Consignee': [consignee],
        'Notify_party': [notify_party],
        'Ocean_vessel': [ocean_vessel],
        'Port_of_loading': [port_of_loading],
        'Port_of_discharge': [port_of_discharge],
        'Place_of_delivery': [place_of_delivery],
        'For_delivery_of_goods': [for_delivery_of_goods],
        'Container_no_and_seal_no': [container_no_and_seal_no],
        'No_of_container': [no_of_container],
        'Gross_weight': [gross_weight],
        'Measurement': [measurement]
    })

# 主流程
pdf = pdfquery.PDFQuery('HBL.PDF')
pagecount = pdf.doc.catalog['Pages'].resolve()['Count']
master = pd.DataFrame()

for p in range(pagecount):
    pdf.load(p)
    page_data = pdfscrape(pdf)
    master = pd.concat([master, page_data], ignore_index=True)

# 导出CSV
master.to_csv('output.csv', index=False)
# 可选:保存XML用于调试
pdf.tree.write('pdfXML.xml', pretty_print=True)

核心改进说明

  • 动态定位逻辑:通过extract_field_by_label函数先找到字段对应的标签文本(如"Shipper"),再根据标签坐标自动定位相邻区域的内容,彻底摆脱硬编码坐标的限制。
  • 适配不同布局:支持right(标签右侧)和below(标签下方)两种方向,通过tolerance参数调整坐标匹配范围,适配不同PDF的布局差异。
  • 多行内容合并:自动合并同一字段的多行文本,避免内容拆分导致的信息缺失。
  • 通用扩展性:只需调整标签文本和方向参数,即可适配不同格式的HBL或其他结构化PDF文件。

内容的提问来源于stack exchange,提问作者Abhinav Bhardwaj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 13:43:04