使用PDFQuery提取单页PDF中重复文本坐标的技术问题
解决方案
针对你遇到的PDFQuery匹配到错误"Result"实例的问题,有两种可行的解决思路:
1. 基于"Other Test Results"的位置限定查找范围
先定位到"Other Test Results"文本的坐标,再只筛选出位于其下方的"Result"实例,代码如下:
from pdfquery import PDFQuery pdf = PDFQuery('example.pdf') pdf.load(0) # 获取"Other Test Results"的顶部y坐标(根据PDF坐标系调整判断逻辑) other_test_header = pdf.pq('LTTextLineHorizontal:contains("Other Test Results")')[0] section_top_y = float(other_test_header.get('y1', 0)) # 提取所有包含"Result"的文本行 all_result_entries = pdf.pq('LTTextLineHorizontal:contains("Result")') # 筛选位于目标区域下方的Result target_result = None for entry in all_result_entries: entry_top_y = float(entry.get('y1', 0)) # 若PDF坐标系中y值向下递增,则改为entry_top_y > section_top_y if entry_top_y < section_top_y: target_result = entry break if target_result: x0 = float(target_result.get('x0', 0)) y0 = float(target_result.get('y0', 0)) x1 = float(target_result.get('x1', 0)) y1 = float(target_result.get('y1', 0)) print(f"目标Result坐标:x0={x0}, y0={y0}, x1={x1}, y1={y1}") else: print("未找到符合条件的Result表头")
2. 结合同组表头的坐标做精准筛选
因为目标"Result"和"Description""Method""Limits"纵向对齐,它们的x坐标范围会高度重叠,可以通过这个特征进一步过滤:
from pdfquery import PDFQuery pdf = PDFQuery('example.pdf') pdf.load(0) # 定位"Other Test Results"区域 other_test_header = pdf.pq('LTTextLineHorizontal:contains("Other Test Results")')[0] section_top_y = float(other_test_header.get('y1', 0)) # 获取同组其他表头的坐标范围 desc = pdf.pq('LTTextLineHorizontal:contains("Description")')[0] method = pdf.pq('LTTextLineHorizontal:contains("Method")')[0] limits = pdf.pq('LTTextLineHorizontal:contains("Limits")')[0] min_x = min(float(desc.get('x0')), float(method.get('x0')), float(limits.get('x0'))) max_x = max(float(desc.get('x1')), float(method.get('x1')), float(limits.get('x1'))) # 筛选符合位置和坐标范围的Result all_result_entries = pdf.pq('LTTextLineHorizontal:contains("Result")') target_result = None for entry in all_result_entries: entry_x0 = float(entry.get('x0')) entry_x1 = float(entry.get('x1')) entry_top_y = float(entry.get('y1')) # 同时满足:在目标区域下方、x坐标与其他表头重叠 if entry_top_y < section_top_y and entry_x0 <= max_x and entry_x1 >= min_x: target_result = entry break if target_result: x0 = float(target_result.get('x0', 0)) y0 = float(target_result.get('y0', 0)) x1 = float(target_result.get('x1', 0)) y1 = float(target_result.get('y1', 0)) print(f"目标Result坐标:x0={x0}, y0={y0}, x1={x1}, y1={y1}") else: print("未找到符合条件的Result表头")
额外提示
- 不同PDF的坐标系可能存在差异,若筛选结果不对,可将
entry_top_y < section_top_y改为entry_top_y > section_top_y尝试; - 若有多个候选"Result",可通过
font或fontsize属性进一步筛选(表头通常使用加粗或更大字号)。
内容的提问来源于stack exchange,提问作者Pd46
相关产品推荐
相关产品推荐

