You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SwiftUI中如何将VNRectangleObservation转换为UIImage?

问题分析与解决方案

你的代码核心问题出在坐标系统的不匹配,Vision框架的boundingBox有两个关键特性你没处理:

  • 是归一化坐标(数值范围0-1),对应图像的相对位置,而非像素坐标
  • 以图像左下角为原点,但CGImage的裁剪坐标是以左上角为原点
    另外你还忽略了UIImage逻辑尺寸和CGImage像素尺寸的scale差异,这些导致裁剪出的图片位置/大小错误。

修正后的裁剪代码

let rectanglesDetection = VNDetectRectanglesRequest { request, error in
    guard let rectangles = request.results as? [VNRectangleObservation], 
          let originalImage = image, 
          let cgImage = originalImage.cgImage else {
        // 处理错误或空结果
        return
    }
    
    // 按视觉从上到下排序(Vision的y轴从下到上,所以取反排序)
    let sortedRectangles = rectangles.sorted { $0.boundingBox.origin.y > $1.boundingBox.origin.y }
    
    let cgImageWidth = cgImage.width
    let cgImageHeight = cgImage.height
    let imageScale = originalImage.scale
    var croppedImages: [UIImage] = []
    
    for rectangle in sortedRectangles {
        // 1. 转换Vision归一化坐标到CGImage像素坐标,同时翻转Y轴
        let x = rectangle.boundingBox.minX * CGFloat(cgImageWidth)
        // Vision的y起点在左下角,CGImage在左上角,所以翻转计算
        let y = CGFloat(cgImageHeight) - (rectangle.boundingBox.minY + rectangle.boundingBox.height) * CGFloat(cgImageHeight)
        let width = rectangle.boundingBox.width * CGFloat(cgImageWidth)
        let height = rectangle.boundingBox.height * CGFloat(cgImageHeight)
        
        let cropRect = CGRect(x: x, y: y, width: width, height: height)
        
        // 2. 裁剪CGImage
        guard let croppedCGImage = cgImage.cropping(to: cropRect) else {
            continue // 裁剪失败则跳过当前矩形
        }
        
        // 3. 转换回UIImage,保留原图像的scale和方向
        let croppedImage = UIImage(cgImage: croppedCGImage, scale: imageScale, orientation: originalImage.imageOrientation)
        croppedImages.append(croppedImage)
    }
    
    self.checkBoxImages = croppedImages
}

针对文字识别的优化建议

你的目标是识别矩形内的文字,不需要提前裁剪图片,直接将VNRectangleObservation设置为文字识别请求的感兴趣区域(ROI)即可,这样更高效:

// 在矩形检测的回调内
for rectangle in sortedRectangles {
    let textRequest = VNRecognizeTextRequest { textRequest, textError in
        guard let textResults = textRequest.results as? [VNRecognizedTextObservation] else { return }
        for result in textResults {
            // 获取置信度最高的识别结果
            guard let topText = result.topCandidates(1).first else { continue }
            print("识别文字:\(topText.string)")
        }
    }
    
    // 设置只识别当前矩形区域
    textRequest.regionOfInterest = rectangle.boundingBox
    
    // 提交文字识别请求
    let handler = VNImageRequestHandler(cgImage: cgImage, orientation: originalImage.imageOrientation)
    try? handler.perform([textRequest])
}

内容的提问来源于stack exchange,提问作者AnujAroshA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 11:40:21