You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SwiftUI中如何通过Vision框架识别文字并获取其位置信息

解决Vision框架OCR无法获取文字位置的问题

你的现有代码仅提取了识别出的文本内容,而VNRecognizedTextObservation对象本身就包含文字的边界框(bounding box)信息,同时每个候选识别结果也支持获取对应位置数据。以下是修改后的代码,可同时获取文本内容与位置信息:

修改后的完整代码

import SwiftUI
import Vision

// 可选:用结构体统一存储文本和位置信息
struct RecognizedTextWithPosition {
    let text: String
    // boundingBox基于图像坐标系(原点在左下角,范围0-1)
    let boundingBox: CGRect
}

struct ContentView: View {
    
    @State var recognizedText = ""
    @State var positionInfo = ""
    
    var body: some View {
        VStack(spacing: 20) {
            Text("OCR using Vision")
                .font(.title)
            
            Image("quote")
                .resizable()
                .scaledToFit()
            
            Button("Recognize Text"){
                ocr()
            }
            
            Text("识别出的文本:")
                .font(.headline)
            TextEditor(text: $recognizedText)
                .frame(height: 150)
                .border(Color.gray)
            
            Text("位置信息:")
                .font(.headline)
            TextEditor(text: $positionInfo)
                .frame(height: 150)
                .border(Color.gray)
        }
        .padding()
        
    }
    
    func ocr() {
        let image = UIImage(named: "quote")
        
        if let cgImage = image?.cgImage {
            
            let handler = VNImageRequestHandler(cgImage: cgImage)
            
            let recognizeRequest = VNRecognizeTextRequest { (request, error) in
                guard let observations = request.results as? [VNRecognizedTextObservation] else {
                    return
                }
                
                var textArray = [String]()
                var positionArray = [String]()
                
                for observation in observations {
                    // 获取匹配度最高的候选文本
                    guard let topCandidate = observation.topCandidates(1).first else { continue }
                    
                    let text = topCandidate.string
                    textArray.append(text)
                    
                    // 获取当前文本块的整体边界框
                    let textBox = observation.boundingBox
                    positionArray.append("文本:\(text)\n位置:\(textBox)")
                    
                    // 若需要字符级位置,可遍历以下内容
                    // for character in topCandidate.characters {
                    //     let charBox = character.boundingBox
                    //     // 处理单字符位置
                    // }
                }
                
                DispatchQueue.main.async {
                    recognizedText = textArray.joined(separator: "\n")
                    positionInfo = positionArray.joined(separator: "\n\n")
                }
            }
            
            recognizeRequest.recognitionLevel = .accurate
            recognizeRequest.usesLanguageCorrection = true
            
            do {
                try handler.perform([recognizeRequest])
            } catch {
                print("识别出错:\(error)")
            }
            
        }
    }
}

关键说明

  • 坐标系转换:Vision返回的boundingBox基于图像坐标系(原点在左下角),如果要适配UIKit的左上角坐标系,可通过以下代码转换:
    guard let image = UIImage(named: "quote") else { return }
    let convertedFrame = CGRect(
        x: textBox.origin.x,
        y: 1 - textBox.origin.y - textBox.size.height,
        width: textBox.size.width,
        height: textBox.size.height
    ).applying(CGAffineTransform(scaleX: image.size.width, y: image.size.height))
    
  • 字符级位置:若需要更精细的单字符位置,可遍历topCandidate.characters,每个VNRecognizedTextCharacter对象都包含boundingBox属性。

内容的提问来源于stack exchange,提问作者Ocwen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 22:45:21