You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在SwiftUI中用Vision框架的VNRecognizeTextRequest(.accurate)提取单词级Bounding Box?

如何用VNRecognizeTextRequest获取单个单词的Bounding Box

要获取单个单词的bounding box,核心是利用VNRecognizedText的segments属性——当使用.accurate识别级别时,Vision会将识别出的文本拆分为单个单词级别的segment,每个segment包含对应单词的文本和独立的bounding box。

关键配置与实现步骤

  • 保持request.recognitionLevel = .accurate以保证识别精度
  • 遍历每个识别结果的候选文本片段(segments),而非直接取整个文本块的bounding box
  • 注意坐标转换:Vision使用基于图像左下角的归一化坐标,需转换为UIKit/SwiftUI的左上角原点坐标系

修改后的完整代码示例

import SwiftUI
import Vision

struct OCR: View {
    @State var image: UIImage? = UIImage(named: "test")
    @State var wordTexts: [String] = []
    @State var wordPositions: [CGRect] = []
    
    @State var size: CGSize = CGSize()
    
    /// 转换Vision归一化坐标到SwiftUI视图坐标
    func convertVisionRectToViewRect(rect: CGRect, imageSize: CGSize) -> CGRect {
        let width = imageSize.width
        let height = imageSize.height
        
        // Vision坐标原点在图像左下角,SwiftUI在左上角,需翻转Y轴
        let x = rect.minX * width
        let y = (1 - rect.maxY) * height
        let rectWidth = rect.width * width
        let rectHeight = rect.height * height
        
        return CGRect(x: x, y: y, width: rectWidth, height: rectHeight)
    }
    
    var body: some View {
        ZStack {
            if let image = image {
                Image(uiImage: image)
                    .resizable()
                    .aspectRatio(contentMode: .fit)
                    .background {
                        GeometryReader { geo in
                            Color.clear
                                .onAppear {
                                    size = geo.size
                                }
                        }
                    }
                    .overlay(Canvas { context, size in
                        for position in wordPositions {
                            let viewRect = convertVisionRectToViewRect(rect: position, imageSize: size)
                            context.stroke(Path(viewRect), with: .color(.red), lineWidth: 1)
                        }
                    })
                    .onAppear {
                        recognizeWordsInImage(image: image) { texts, positions in
                            wordTexts = texts
                            wordPositions = positions
                        }
                    }
            } else {
                Text("没有图片")
            }
        }
    }
}

extension OCR {
    func recognizeWordsInImage(image: UIImage, completion: @escaping([String], [CGRect]) -> Void) {
        var words: [String] = []
        var wordBoxes: [CGRect] = []
        
        guard let cgImage = image.cgImage else { 
            completion([], [])
            return 
        }
        
        let request = VNRecognizeTextRequest { (request, error) in
            guard let observations = request.results as? [VNRecognizedTextObservation], error == nil else {
                print("文字识别错误: \(error?.localizedDescription ?? "未知错误")")
                completion([], [])
                return
            }
            
            for observation in observations {
                guard let topCandidate = observation.topCandidates(1).first else { continue }
                // 遍历每个单词segment
                for segment in topCandidate.segments {
                    words.append(segment.string)
                    wordBoxes.append(segment.boundingBox)
                }
            }
            
            DispatchQueue.main.async {
                completion(words, wordBoxes)
            }
        }
        
        // 关键配置:高精度识别
        request.recognitionLevel = .accurate
        // 可选:如果不需要自动纠错,可以关闭(不影响单词分割,但可能改变识别结果)
        // request.usesLanguageCorrection = false
        
        let handler = VNImageRequestHandler(cgImage: cgImage)
        try? handler.perform([request])
    }
}

#Preview {
    OCR()
}

注意事项

  • .accurate模式是获取单词级segment的前提,.fast模式不支持文本分段
  • 部分场景下,Vision可能会将连写词(如带连字符的单词)识别为单个segment,这是正常的识别逻辑
  • 坐标转换函数需根据实际视图的显示模式(如.fit或.fill)调整,确保bounding box位置准确

内容的提问来源于stack exchange,提问作者J W

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 06:03:21