You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用C#实现Google Vision API OCR结果按行拆分?

解决Google Vision API多行文本识别按行拆分的问题(C#)

嗨,我之前也踩过这个坑!一开始只盯着API返回的那串合并文本,后来才发现Google Vision其实藏了更有用的布局信息,给你两个靠谱的解决思路:

方案一:利用API原生的布局分析(首选)

你大概率用的是普通的TextDetection接口,它确实会把所有文本合并成一个字符串,但Google Vision还有个**文档文本检测(Document Text Detection)**接口,专门针对排版复杂的文本,会返回详细的层级结构——包括每行的边界框、单词甚至单个字符的位置。

在C#里,你可以这样解析响应来提取每行文本:

using Google.Cloud.Vision.V1;
using System.Linq;
using System.Collections.Generic;
using System.Threading.Tasks;

public async Task<List<string>> ExtractTextLinesFromImage(string imageFilePath)
{
    // 初始化Vision客户端
    var visionClient = ImageAnnotatorClient.Create();
    var targetImage = Image.FromFile(imageFilePath);

    // 调用文档文本检测接口,这是关键!
    var detectionResponse = await visionClient.DetectDocumentTextAsync(targetImage);

    var extractedLines = new List<string>();

    // 遍历返回的页面、区块、段落(每个段落对应一行文本)
    foreach (var page in detectionResponse.Pages)
    {
        foreach (var block in page.Blocks)
        {
            foreach (var paragraph in block.Paragraphs)
            {
                // 把段落里的所有单词拼接成完整的行文本
                string lineContent = string.Join(" ", paragraph.Words.Select(word => 
                    string.Join("", word.Symbols.Select(symbol => symbol.Text))));
                extractedLines.Add(lineContent);
            }
        }
    }

    return extractedLines;
}

这个方法完全依赖API的专业排版分析,不管你的行是全大写还是混合大小写,都能精准拆分,比自己写字符串过滤逻辑靠谱多了。

方案二:固定位置裁剪图片(适合行坐标完全不变的场景)

如果你的图片每行的位置、尺寸完全固定,那直接裁剪每行区域再单独识别也是个简单粗暴的好办法。在C#里可以用System.Drawing.Common或者更现代的SixLabors.ImageSharp来处理图片裁剪。

用System.Drawing实现的示例(需要先安装System.Drawing.Common NuGet包)

using System.Drawing;
using System.IO;
using Google.Cloud.Vision.V1;
using System.Collections.Generic;
using System.Threading.Tasks;

// 裁剪指定区域的图片
private Image CropImageToRegion(Image originalImage, Rectangle cropArea)
{
    var croppedBitmap = new Bitmap(cropArea.Width, cropArea.Height);
    using (var graphics = Graphics.FromImage(croppedBitmap))
    {
        // 绘制裁剪后的区域
        graphics.DrawImage(originalImage, 
            new Rectangle(0, 0, croppedBitmap.Width, croppedBitmap.Height),
            cropArea,
            GraphicsUnit.Pixel);
    }
    return croppedBitmap;
}

// 按固定区域拆分并识别每行文本
public async Task<List<string>> GetLinesByFixedCrop(string imageFilePath)
{
    var lineTexts = new List<string>();
    var originalImage = Image.FromFile(imageFilePath);
    var visionClient = ImageAnnotatorClient.Create();

    // 这里替换成你实际测量的每行坐标(x, y, 宽度, 高度)
    var lineRegions = new List<Rectangle>
    {
        new Rectangle(60, 120, 780, 45), // 第一行区域
        new Rectangle(60, 200, 780, 45), // 第二行区域
        // 继续添加其他行的区域...
    };

    foreach (var region in lineRegions)
    {
        using (var croppedImage = CropImageToRegion(originalImage, region))
        {
            // 将裁剪后的图片转成MemoryStream供Vision API读取
            using (var imageStream = new MemoryStream())
            {
                croppedImage.Save(imageStream, ImageFormat.Png);
                imageStream.Position = 0;
                var targetImage = Image.FromStream(imageStream);

                // 调用普通文本检测接口即可,因为单区域只有一行
                var detectionResult = await visionClient.DetectTextAsync(targetImage);
                if (detectionResult.Any())
                {
                    lineTexts.Add(detectionResult[0].Description);
                }
            }
        }
    }

    originalImage.Dispose();
    return lineTexts;
}

要获取准确的裁剪坐标,你可以用Photoshop、GIMP或者免费的截图工具(比如Snipaste)来测量每行的位置和尺寸,只要图片分辨率不变,这些坐标就能一直复用。

总结

优先选方案一,它更灵活,哪怕行位置有微小偏移也能正确识别;如果你的图片排版完全固定,方案二会更省心。

内容的提问来源于stack exchange,提问作者Nate

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:22:28