如何用C#实现Google Vision API OCR结果按行拆分?
解决Google Vision API多行文本识别按行拆分的问题(C#)
嗨,我之前也踩过这个坑!一开始只盯着API返回的那串合并文本,后来才发现Google Vision其实藏了更有用的布局信息,给你两个靠谱的解决思路:
方案一:利用API原生的布局分析(首选)
你大概率用的是普通的TextDetection接口,它确实会把所有文本合并成一个字符串,但Google Vision还有个**文档文本检测(Document Text Detection)**接口,专门针对排版复杂的文本,会返回详细的层级结构——包括每行的边界框、单词甚至单个字符的位置。
在C#里,你可以这样解析响应来提取每行文本:
using Google.Cloud.Vision.V1; using System.Linq; using System.Collections.Generic; using System.Threading.Tasks; public async Task<List<string>> ExtractTextLinesFromImage(string imageFilePath) { // 初始化Vision客户端 var visionClient = ImageAnnotatorClient.Create(); var targetImage = Image.FromFile(imageFilePath); // 调用文档文本检测接口,这是关键! var detectionResponse = await visionClient.DetectDocumentTextAsync(targetImage); var extractedLines = new List<string>(); // 遍历返回的页面、区块、段落(每个段落对应一行文本) foreach (var page in detectionResponse.Pages) { foreach (var block in page.Blocks) { foreach (var paragraph in block.Paragraphs) { // 把段落里的所有单词拼接成完整的行文本 string lineContent = string.Join(" ", paragraph.Words.Select(word => string.Join("", word.Symbols.Select(symbol => symbol.Text)))); extractedLines.Add(lineContent); } } } return extractedLines; }
这个方法完全依赖API的专业排版分析,不管你的行是全大写还是混合大小写,都能精准拆分,比自己写字符串过滤逻辑靠谱多了。
方案二:固定位置裁剪图片(适合行坐标完全不变的场景)
如果你的图片每行的位置、尺寸完全固定,那直接裁剪每行区域再单独识别也是个简单粗暴的好办法。在C#里可以用System.Drawing.Common或者更现代的SixLabors.ImageSharp来处理图片裁剪。
用System.Drawing实现的示例(需要先安装System.Drawing.Common NuGet包)
using System.Drawing; using System.IO; using Google.Cloud.Vision.V1; using System.Collections.Generic; using System.Threading.Tasks; // 裁剪指定区域的图片 private Image CropImageToRegion(Image originalImage, Rectangle cropArea) { var croppedBitmap = new Bitmap(cropArea.Width, cropArea.Height); using (var graphics = Graphics.FromImage(croppedBitmap)) { // 绘制裁剪后的区域 graphics.DrawImage(originalImage, new Rectangle(0, 0, croppedBitmap.Width, croppedBitmap.Height), cropArea, GraphicsUnit.Pixel); } return croppedBitmap; } // 按固定区域拆分并识别每行文本 public async Task<List<string>> GetLinesByFixedCrop(string imageFilePath) { var lineTexts = new List<string>(); var originalImage = Image.FromFile(imageFilePath); var visionClient = ImageAnnotatorClient.Create(); // 这里替换成你实际测量的每行坐标(x, y, 宽度, 高度) var lineRegions = new List<Rectangle> { new Rectangle(60, 120, 780, 45), // 第一行区域 new Rectangle(60, 200, 780, 45), // 第二行区域 // 继续添加其他行的区域... }; foreach (var region in lineRegions) { using (var croppedImage = CropImageToRegion(originalImage, region)) { // 将裁剪后的图片转成MemoryStream供Vision API读取 using (var imageStream = new MemoryStream()) { croppedImage.Save(imageStream, ImageFormat.Png); imageStream.Position = 0; var targetImage = Image.FromStream(imageStream); // 调用普通文本检测接口即可,因为单区域只有一行 var detectionResult = await visionClient.DetectTextAsync(targetImage); if (detectionResult.Any()) { lineTexts.Add(detectionResult[0].Description); } } } } originalImage.Dispose(); return lineTexts; }
要获取准确的裁剪坐标,你可以用Photoshop、GIMP或者免费的截图工具(比如Snipaste)来测量每行的位置和尺寸,只要图片分辨率不变,这些坐标就能一直复用。
总结
优先选方案一,它更灵活,哪怕行位置有微小偏移也能正确识别;如果你的图片排版完全固定,方案二会更省心。
内容的提问来源于stack exchange,提问作者Nate
相关产品推荐
相关产品推荐

