You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyMuPDF提取QuadPoints适配Adobe Embed API显示异常问题

解决PyMuPDF提取批注转Adobe Embed API显示异常问题

核心问题分析

  • 坐标系统不匹配:PyMuPDF以页面左下角为坐标原点,Adobe Embed API则以左上角为原点,直接复用PyMuPDF的Y坐标会导致批注位置偏移。
  • QuadPoints格式错误:Adobe高亮批注要求quadPoints为每个高亮块的四个顶点(左上、右上、右下、左下)数组,原代码将四边形转为矩形对角坐标,丢失了精确形状信息,导致显示变形。
  • 缺失Target Source字段:Adobe API要求Annotation的Target必须包含source字段指向目标PDF标识,原代码缺失该字段可能导致定位失败。

修复后的代码

import fitz
import json
import sys

if len(sys.argv) != 2:
    print('Usage: python extractPDFAnnotations.py <filename>')
    sys.exit(1)
filename = sys.argv[1]

doc = fitz.open('Sandborn2003Annotated.pdf')
annotations = []

for page_num, page in enumerate(doc):
    page_rect = page.rect
    page_height = page_rect.height

    for annot in page.annots():
        annotation_data = {}
        target_data = {}
        selector_data = {}

        # 通用批注属性
        annotation_data["@context"] = [
            "https://www.w3.org/ns/anno.jsonld",
            "https://comments.acrobat.com/ns/anno.jsonld",
        ]
        annotation_data["id"] = annot.info.get("id", "")
        annotation_data["type"] = "Annotation"
        annotation_data["motivation"] = "commenting"
        annotation_data["bodyValue"] = annot.info.get("content", "")

        # Target属性:补充source字段(替换为目标PDF的标识/路径)
        target_data["source"] = "Sandborn2003.pdf"
        target_data["selector"] = selector_data
        annotation_data["target"] = target_data

        # Selector基础属性
        selector_data["node"] = {"index": page_num}
        selector_data["type"] = "AdobeAnnoSelector"

        if annot.type[0] == 8:  # 处理高亮批注
            all_coordinates = annot.vertices
            quad_points = []

            # 转换每个顶点的Y坐标,保留原始四边形结构
            for i in range(0, len(all_coordinates), 4):
                quad = all_coordinates[i:i + 4]
                converted_quad = [
                    quad[0].x, page_height - quad[0].y,
                    quad[1].x, page_height - quad[1].y,
                    quad[2].x, page_height - quad[2].y,
                    quad[3].x, page_height - quad[3].y
                ]
                quad_points.extend(converted_quad)

            selector_data["quadPoints"] = quad_points
            selector_data["opacity"] = 0.4
            selector_data["subtype"] = "highlight"
            
            # 转换boundingBox的Y坐标(注意PyMuPDF与Adobe的Y轴方向差异)
            selector_data["boundingBox"] = [
                annot.rect.x0,
                page_height - annot.rect.y1,
                annot.rect.x1,
                page_height - annot.rect.y0,
            ]
            selector_data["strokeColor"] = "#fccb00"
            selector_data["strokeWidth"] = 3

        # Creator属性
        annotation_data["creator"] = {
            "name": annot.info.get("title", ""),
            "type": "Person",
        }

        annotations.append(annotation_data)

# 保存转换后的批注JSON
with open(filename, 'w') as f:
    json.dump(annotations, f, indent=4)

print(f'Annotations saved to {filename}')

关键修复说明

  1. 坐标系统转换:通过page_height - y将PyMuPDF的Y坐标转换为Adobe标准,同时调整boundingBox的Y值顺序(PyMuPDF矩形的y0为底部,y1为顶部,Adobe则相反)。
  2. QuadPoints格式修正:直接转换原始四边形的每个顶点坐标,保留高亮的精确形状,确保每个高亮块的quadPoints包含8个数值(4个顶点×2坐标)。
  3. 补充Source字段:在Target中添加source字段,明确指定批注要应用的目标PDF,确保Adobe API能正确定位。

验证步骤

  • 检查生成的JSON中quadPoints的每个高亮块是否为8个数值的数组。
  • 确认boundingBox的Y值范围符合目标PDF的页面高度(例如A4页面高度为842,转换后Y值应在0-842之间)。
  • 确保Adobe Embed API调用时使用的PDF文件与source字段的标识完全匹配。

内容的提问来源于stack exchange,提问作者Justin Erswell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 08:35:02