PyMuPDF提取QuadPoints适配Adobe Embed API显示异常问题
解决PyMuPDF提取批注转Adobe Embed API显示异常问题
核心问题分析
- 坐标系统不匹配:PyMuPDF以页面左下角为坐标原点,Adobe Embed API则以左上角为原点,直接复用PyMuPDF的Y坐标会导致批注位置偏移。
- QuadPoints格式错误:Adobe高亮批注要求
quadPoints为每个高亮块的四个顶点(左上、右上、右下、左下)数组,原代码将四边形转为矩形对角坐标,丢失了精确形状信息,导致显示变形。 - 缺失Target Source字段:Adobe API要求Annotation的Target必须包含
source字段指向目标PDF标识,原代码缺失该字段可能导致定位失败。
修复后的代码
import fitz import json import sys if len(sys.argv) != 2: print('Usage: python extractPDFAnnotations.py <filename>') sys.exit(1) filename = sys.argv[1] doc = fitz.open('Sandborn2003Annotated.pdf') annotations = [] for page_num, page in enumerate(doc): page_rect = page.rect page_height = page_rect.height for annot in page.annots(): annotation_data = {} target_data = {} selector_data = {} # 通用批注属性 annotation_data["@context"] = [ "https://www.w3.org/ns/anno.jsonld", "https://comments.acrobat.com/ns/anno.jsonld", ] annotation_data["id"] = annot.info.get("id", "") annotation_data["type"] = "Annotation" annotation_data["motivation"] = "commenting" annotation_data["bodyValue"] = annot.info.get("content", "") # Target属性:补充source字段(替换为目标PDF的标识/路径) target_data["source"] = "Sandborn2003.pdf" target_data["selector"] = selector_data annotation_data["target"] = target_data # Selector基础属性 selector_data["node"] = {"index": page_num} selector_data["type"] = "AdobeAnnoSelector" if annot.type[0] == 8: # 处理高亮批注 all_coordinates = annot.vertices quad_points = [] # 转换每个顶点的Y坐标,保留原始四边形结构 for i in range(0, len(all_coordinates), 4): quad = all_coordinates[i:i + 4] converted_quad = [ quad[0].x, page_height - quad[0].y, quad[1].x, page_height - quad[1].y, quad[2].x, page_height - quad[2].y, quad[3].x, page_height - quad[3].y ] quad_points.extend(converted_quad) selector_data["quadPoints"] = quad_points selector_data["opacity"] = 0.4 selector_data["subtype"] = "highlight" # 转换boundingBox的Y坐标(注意PyMuPDF与Adobe的Y轴方向差异) selector_data["boundingBox"] = [ annot.rect.x0, page_height - annot.rect.y1, annot.rect.x1, page_height - annot.rect.y0, ] selector_data["strokeColor"] = "#fccb00" selector_data["strokeWidth"] = 3 # Creator属性 annotation_data["creator"] = { "name": annot.info.get("title", ""), "type": "Person", } annotations.append(annotation_data) # 保存转换后的批注JSON with open(filename, 'w') as f: json.dump(annotations, f, indent=4) print(f'Annotations saved to {filename}')
关键修复说明
- 坐标系统转换:通过
page_height - y将PyMuPDF的Y坐标转换为Adobe标准,同时调整boundingBox的Y值顺序(PyMuPDF矩形的y0为底部,y1为顶部,Adobe则相反)。 - QuadPoints格式修正:直接转换原始四边形的每个顶点坐标,保留高亮的精确形状,确保每个高亮块的
quadPoints包含8个数值(4个顶点×2坐标)。 - 补充Source字段:在Target中添加
source字段,明确指定批注要应用的目标PDF,确保Adobe API能正确定位。
验证步骤
- 检查生成的JSON中
quadPoints的每个高亮块是否为8个数值的数组。 - 确认
boundingBox的Y值范围符合目标PDF的页面高度(例如A4页面高度为842,转换后Y值应在0-842之间)。 - 确保Adobe Embed API调用时使用的PDF文件与
source字段的标识完全匹配。
内容的提问来源于stack exchange,提问作者Justin Erswell
相关产品推荐
相关产品推荐

