如何将Document AI的归一化坐标转换为文档原始尺度并裁剪图像?
解决Google Cloud Document AI归一化坐标转像素并裁剪PDF页面的方案
步骤1:提取关键数据
从Document AI的响应中获取两个核心信息:
- 实体的归一化边界框坐标:
result.document.entities[0].page_anchor.page_refs[0].bounding_poly.normalized_vertices - 页面的像素尺寸:
result.document.pages[0].dimension(包含width、height字段)
步骤2:归一化坐标转像素坐标
归一化坐标是0-1的比例值,直接乘以对应页面的宽高即可得到像素坐标。边界框顶点顺序为左上、右上、右下、左下,只需取左上和右下顶点就能确定裁剪区域。
示例代码:
# 从Document AI响应提取数据 normalized_vertices = result.document.entities[0].page_anchor.page_refs[0].bounding_poly.normalized_vertices page_dim = result.document.pages[0].dimension # 获取页面像素宽高 page_width = page_dim.width page_height = page_dim.height # 转换左上、右下顶点为像素坐标 x1 = int(normalized_vertices[0].x * page_width) y1 = int(normalized_vertices[0].y * page_height) x2 = int(normalized_vertices[2].x * page_width) y2 = int(normalized_vertices[2].y * page_height)
步骤3:PDF转图像并裁剪
用pdf2image将PDF转为图像,再通过cv2完成区域裁剪:
from pdf2image import convert_from_path import cv2 import numpy as np # 转换PDF第一页为PIL图像(与Document AI处理页面对应) pdf_pages = convert_from_path("target_document.pdf") pil_image = pdf_pages[0] # 转换为cv2兼容的BGR格式 cv2_image = cv2.cvtColor(np.array(pil_image), cv2.COLOR_RGB2BGR) # 执行裁剪 cropped_image = cv2_image[y1:y2, x1:x2] # 保存或预览结果 cv2.imwrite("cropped_entity.png", cropped_image) cv2.imshow("Cropped Result", cropped_image) cv2.waitKey(0) cv2.destroyAllWindows()
注意事项
- 确保
pdf2image转换图像的分辨率与Document AI识别的页面尺寸匹配,若转换时指定了自定义dpi,需按比例调整坐标计算逻辑。 - 若裁剪区域偏移,可检查顶点索引是否正确(比如验证4个顶点的坐标值,确认边界框范围)。
内容的提问来源于stack exchange,提问作者Rajiv2806
相关产品推荐
相关产品推荐

