单页PDF转UIImage后文本识别质量不佳的优化方案求助
单页PDF转UIImage后文本识别质量不佳的优化方案求助
各位大佬好,这个问题我卡了好久了,实在没辙了来求帮忙!
我现在要处理用户的PDF格式每日任务文档,核心是要从里面提取最多20组「日期+机场/时间」的组合。原本直接提取PDF文本的时候,内容的顺序完全乱套,日期和对应的机场/时间根本对不上,试了.string和.attributedString两种方式都解决不了顺序问题。
后来我改成了把PDF转成UIImage,再用Vision框架提取带排版结构的文本——这个思路逻辑上是对的,能保证内容顺序和原PDF一致,但新问题来了:PDF转图片的过程中质量损失太严重,要么漏识别文本,要么直接识别错误,完全没法稳定工作。肯定有更稳健的PDF转图+文本提取的方法吧?
我先贴一下目前用的PDF转UIImage的代码:
func drawPDFfromURL(url: URL) -> UIImage? { guard let document = PDFDocument(url: url), let page = document.page(at: 0) else { return nil } let scale = CGFloat(3.0) let pageRect = page.bounds(for: .mediaBox) let scaledSize = CGSize(width: pageRect.width * scale, height: pageRect.height * scale) let rendererFormat = UIGraphicsImageRendererFormat() rendererFormat.scale = 1 // 也试过用UIScreen.main.scale rendererFormat.opaque = false let renderer = UIGraphicsImageRenderer(size: scaledSize, format: rendererFormat) let image = renderer.image { ctx in let context = ctx.cgContext UIColor.lightText.set() // 设成lightText解决了不少问题 context.fill(CGRect(origin: .zero, size: scaledSize)) context.saveGState() // 翻转和缩放以正确渲染 context.translateBy(x: 0.0, y: scaledSize.height) context.scaleBy(x: scale, y: -scale) // 尝试增强细字体的线条宽度 context.setLineWidth(2.0) context.setLineJoin(.round) context.setLineCap(.round) page.draw(with: .mediaBox, to: context) context.restoreGState() } return image }
再贴一下Vision文本提取的代码:
func getTextFromImage(image: UIImage) { guard let cgImage = image.cgImage else { return } let request = VNRecognizeTextRequest { request, error in guard let results = request.results as? [VNRecognizedTextObservation] else { print("No text found.") return } var wordItems: [(text: String, rect: CGRect)] = [] for observation in results { guard let candidate = observation.topCandidates(1).first else { continue } let stringRange = candidate.string.startIndex..<candidate.string.endIndex if let boxObservation = try? candidate.boundingBox(for: stringRange) { let normBox = boxObservation.boundingBox let pixelBox = VNImageRectForNormalizedRect( normBox, Int(image.size.width), Int(image.size.height) ) wordItems.append((candidate.string, pixelBox)) } } // 按Y坐标分组(行逻辑),Y容差控制每行的匹配范围 let yTolerance: CGFloat = 35.0 let grouped = Dictionary(grouping: wordItems) { word in round(word.rect.origin.y / yTolerance) * yTolerance } // 从上到下排序行 let sortedLines = grouped.sorted(by: { $0.key > $1.key }) // 每行内从左到右排序单词并拼接 let lines = sortedLines.map { (_, words) in words.sorted(by: { $0.rect.origin.x < $1.rect.origin.x }) .map(\.text) .joined(separator: " ") } let rawText = lines.joined(separator: "\n") print("OCR with Layout:\n\(rawText)") } request.recognitionLevel = .accurate request.recognitionLanguages = ["en-US"] let handler = VNImageRequestHandler(cgImage: cgImage, options: [:]) try? handler.perform([request]) }
我自己也试过调整不少参数:比如UIColor.lightText.set()确实解决了不少显示问题,也试过改缩放比例、yTolerance、setLineWidth这些,但始终没法解决图片质量导致的识别误差问题。有没有更靠谱的PDF转图+文本提取的方案?求各位大佬指点迷津!
内容来源于stack exchange
相关产品推荐
相关产品推荐

