You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PDFQuery提取单页PDF中重复文本坐标的技术问题

解决方案

针对你遇到的PDFQuery匹配到错误"Result"实例的问题,有两种可行的解决思路:

1. 基于"Other Test Results"的位置限定查找范围

先定位到"Other Test Results"文本的坐标,再只筛选出位于其下方的"Result"实例,代码如下:

from pdfquery import PDFQuery

pdf = PDFQuery('example.pdf')
pdf.load(0)

# 获取"Other Test Results"的顶部y坐标(根据PDF坐标系调整判断逻辑)
other_test_header = pdf.pq('LTTextLineHorizontal:contains("Other Test Results")')[0]
section_top_y = float(other_test_header.get('y1', 0))

# 提取所有包含"Result"的文本行
all_result_entries = pdf.pq('LTTextLineHorizontal:contains("Result")')

# 筛选位于目标区域下方的Result
target_result = None
for entry in all_result_entries:
    entry_top_y = float(entry.get('y1', 0))
    # 若PDF坐标系中y值向下递增,则改为entry_top_y > section_top_y
    if entry_top_y < section_top_y:
        target_result = entry
        break

if target_result:
    x0 = float(target_result.get('x0', 0))
    y0 = float(target_result.get('y0', 0))
    x1 = float(target_result.get('x1', 0))
    y1 = float(target_result.get('y1', 0))
    print(f"目标Result坐标:x0={x0}, y0={y0}, x1={x1}, y1={y1}")
else:
    print("未找到符合条件的Result表头")

2. 结合同组表头的坐标做精准筛选

因为目标"Result"和"Description""Method""Limits"纵向对齐,它们的x坐标范围会高度重叠,可以通过这个特征进一步过滤:

from pdfquery import PDFQuery

pdf = PDFQuery('example.pdf')
pdf.load(0)

# 定位"Other Test Results"区域
other_test_header = pdf.pq('LTTextLineHorizontal:contains("Other Test Results")')[0]
section_top_y = float(other_test_header.get('y1', 0))

# 获取同组其他表头的坐标范围
desc = pdf.pq('LTTextLineHorizontal:contains("Description")')[0]
method = pdf.pq('LTTextLineHorizontal:contains("Method")')[0]
limits = pdf.pq('LTTextLineHorizontal:contains("Limits")')[0]

min_x = min(float(desc.get('x0')), float(method.get('x0')), float(limits.get('x0')))
max_x = max(float(desc.get('x1')), float(method.get('x1')), float(limits.get('x1')))

# 筛选符合位置和坐标范围的Result
all_result_entries = pdf.pq('LTTextLineHorizontal:contains("Result")')
target_result = None

for entry in all_result_entries:
    entry_x0 = float(entry.get('x0'))
    entry_x1 = float(entry.get('x1'))
    entry_top_y = float(entry.get('y1'))
    
    # 同时满足:在目标区域下方、x坐标与其他表头重叠
    if entry_top_y < section_top_y and entry_x0 <= max_x and entry_x1 >= min_x:
        target_result = entry
        break

if target_result:
    x0 = float(target_result.get('x0', 0))
    y0 = float(target_result.get('y0', 0))
    x1 = float(target_result.get('x1', 0))
    y1 = float(target_result.get('y1', 0))
    print(f"目标Result坐标:x0={x0}, y0={y0}, x1={x1}, y1={y1}")
else:
    print("未找到符合条件的Result表头")

额外提示

  • 不同PDF的坐标系可能存在差异,若筛选结果不对,可将entry_top_y < section_top_y改为entry_top_y > section_top_y尝试;
  • 若有多个候选"Result",可通过font或fontsize属性进一步筛选(表头通常使用加粗或更大字号)。

内容的提问来源于stack exchange,提问作者Pd46

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 13:37:49