You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Document AI的归一化坐标转换为文档原始尺度并裁剪图像?

解决Google Cloud Document AI归一化坐标转像素并裁剪PDF页面的方案

步骤1:提取关键数据

从Document AI的响应中获取两个核心信息:

  • 实体的归一化边界框坐标:result.document.entities[0].page_anchor.page_refs[0].bounding_poly.normalized_vertices
  • 页面的像素尺寸:result.document.pages[0].dimension(包含width、height字段)

步骤2:归一化坐标转像素坐标

归一化坐标是0-1的比例值,直接乘以对应页面的宽高即可得到像素坐标。边界框顶点顺序为左上、右上、右下、左下,只需取左上和右下顶点就能确定裁剪区域。

示例代码:

# 从Document AI响应提取数据
normalized_vertices = result.document.entities[0].page_anchor.page_refs[0].bounding_poly.normalized_vertices
page_dim = result.document.pages[0].dimension

# 获取页面像素宽高
page_width = page_dim.width
page_height = page_dim.height

# 转换左上、右下顶点为像素坐标
x1 = int(normalized_vertices[0].x * page_width)
y1 = int(normalized_vertices[0].y * page_height)
x2 = int(normalized_vertices[2].x * page_width)
y2 = int(normalized_vertices[2].y * page_height)

步骤3:PDF转图像并裁剪

用pdf2image将PDF转为图像,再通过cv2完成区域裁剪:

from pdf2image import convert_from_path
import cv2
import numpy as np

# 转换PDF第一页为PIL图像(与Document AI处理页面对应)
pdf_pages = convert_from_path("target_document.pdf")
pil_image = pdf_pages[0]

# 转换为cv2兼容的BGR格式
cv2_image = cv2.cvtColor(np.array(pil_image), cv2.COLOR_RGB2BGR)

# 执行裁剪
cropped_image = cv2_image[y1:y2, x1:x2]

# 保存或预览结果
cv2.imwrite("cropped_entity.png", cropped_image)
cv2.imshow("Cropped Result", cropped_image)
cv2.waitKey(0)
cv2.destroyAllWindows()

注意事项

  • 确保pdf2image转换图像的分辨率与Document AI识别的页面尺寸匹配,若转换时指定了自定义dpi,需按比例调整坐标计算逻辑。
  • 若裁剪区域偏移,可检查顶点索引是否正确(比如验证4个顶点的坐标值,确认边界框范围)。

内容的提问来源于stack exchange,提问作者Rajiv2806

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 06:00:32