You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在C#中使用Google Vision API提取PDF的文本与表格内容

C# 调用Google Vision API提取PDF中表格的实现方案

问题原因说明

你之前运行报错是因为DetectDocumentText仅支持JPG/PNG等图片格式输入,无法直接处理PDF文件,PDF这类多页文档的结构化内容提取必须使用异步文件批处理接口。

前置准备

  • 确认已安装最新版NuGet包:Google.Cloud.Vision.V1(旧版Google.Cloud.Vision.API已停止维护,建议替换)
  • 待处理PDF优先存储到Google Cloud Storage(GCS),小体积PDF也可直接使用本地字节流
  • 你的服务账号JSON密钥有对应GCS桶的读写权限(如果使用GCS存储的话)

实现步骤

步骤1:提交PDF文档异步检测请求

using Google.Cloud.Vision.V1;
using Google.LongRunning;

// 初始化客户端,传入你的服务账号密钥路径
var client = await new ImageAnnotatorClientBuilder
{
    CredentialsPath = @"myjsonfile.json"
}.BuildAsync();

// 配置输入源,以下二选一即可
// 选项1:读取GCS存储的PDF
var inputConfig = new InputConfig
{
    GcsSource = new GcsSource { Uri = "gs://你的GCS桶名/待处理文件.pdf" },
    MimeType = "application/pdf"
};
// 选项2:读取本地PDF(仅适合小体积文件)
// var inputConfig = new InputConfig
// {
//     Content = Google.Protobuf.ByteString.FromStream(File.OpenRead(@"本地PDF路径.pdf")),
//     MimeType = "application/pdf"
// };

// 配置检测类型为文档文本检测
var feature = new Feature { Type = Feature.Types.Type.DocumentTextDetection };

// 配置结果输出路径(必须为GCS路径,小文件后续也可以直接读取结果无需下载)
var outputConfig = new OutputConfig
{
    GcsDestination = new GcsDestination { Uri = "gs://你的GCS桶名/结果输出前缀/" },
    // 单个输出JSON最多包含的页数,最大支持100
    BatchSize = 10
};

// 构造并提交异步处理请求
var request = new AsyncAnnotateFileRequest
{
    InputConfig = inputConfig,
    Features = { feature },
    OutputConfig = outputConfig
};
Operation<AsyncBatchAnnotateFilesResponse, OperationMetadata> operation = 
    await client.AsyncBatchAnnotateFilesAsync(new[] { request });

// 等待作业完成,大文件可以自行实现轮询逻辑替代阻塞等待
await operation.PollUntilCompletedAsync();

步骤2:解析检测结果提取表格内容

作业完成后,你指定的GCS输出路径下会生成output-1-to-10.json这类结构的结果文件,下载后按以下逻辑解析即可:

using Google.Cloud.Vision.V1;
using Google.Protobuf;

// 读取结果JSON
var resultJson = File.ReadAllText(@"下载的结果JSON路径.json");
var batchResponse = BatchAnnotateFilesResponse.Parser.ParseJson(resultJson);

// 遍历所有页面的标注结果
foreach (var fileResp in batchResponse.Responses)
{
    var fullText = fileResp.FullTextAnnotation;
    // 筛选所有表格类型的Block
    var tableBlocks = fullText.Pages
        .SelectMany(p => p.Blocks)
        .Where(b => b.BlockType == Block.Types.BlockType.Table);
    
    foreach (var tableBlock in tableBlocks)
    {
        var table = tableBlock.Table;
        Console.WriteLine($"匹配到表格:{table.HeaderRows.Count}行表头,{table.BodyRows.Count}行表体");
        // 拼接表格内容,可自行调整输出格式
        foreach (var row in table.HeaderRows.Concat(table.BodyRows))
        {
            var rowCells = row.Cells
                .Select(cell => 
                    // 直接取单元格文本,如需保留格式可以用fullText.Text.Substring(cell.TextOffset, cell.TextLength)
                    string.Join("", cell.Words.SelectMany(w => w.Symbols).Select(s => s.Text)).Trim()
                );
            Console.WriteLine(string.Join("\t", rowCells));
        }
    }
}

注意事项

  • 单份PDF大小不能超过2GB,页数上限为2000页
  • 异步作业执行时间和PDF页数正相关,大文件建议加超时和异常重试逻辑
  • 如需提取表格的合并单元格、边框等属性,可以直接读取Table、TableRow、TableCell对象的对应字段获取

内容的提问来源于stack exchange,提问作者Kristina Hammer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 02:06:05