如何在C#中使用Google Vision API提取PDF的文本与表格内容
C# 调用Google Vision API提取PDF中表格的实现方案
问题原因说明
你之前运行报错是因为DetectDocumentText仅支持JPG/PNG等图片格式输入,无法直接处理PDF文件,PDF这类多页文档的结构化内容提取必须使用异步文件批处理接口。
前置准备
- 确认已安装最新版NuGet包:
Google.Cloud.Vision.V1(旧版Google.Cloud.Vision.API已停止维护,建议替换) - 待处理PDF优先存储到Google Cloud Storage(GCS),小体积PDF也可直接使用本地字节流
- 你的服务账号JSON密钥有对应GCS桶的读写权限(如果使用GCS存储的话)
实现步骤
步骤1:提交PDF文档异步检测请求
using Google.Cloud.Vision.V1; using Google.LongRunning; // 初始化客户端,传入你的服务账号密钥路径 var client = await new ImageAnnotatorClientBuilder { CredentialsPath = @"myjsonfile.json" }.BuildAsync(); // 配置输入源,以下二选一即可 // 选项1:读取GCS存储的PDF var inputConfig = new InputConfig { GcsSource = new GcsSource { Uri = "gs://你的GCS桶名/待处理文件.pdf" }, MimeType = "application/pdf" }; // 选项2:读取本地PDF(仅适合小体积文件) // var inputConfig = new InputConfig // { // Content = Google.Protobuf.ByteString.FromStream(File.OpenRead(@"本地PDF路径.pdf")), // MimeType = "application/pdf" // }; // 配置检测类型为文档文本检测 var feature = new Feature { Type = Feature.Types.Type.DocumentTextDetection }; // 配置结果输出路径(必须为GCS路径,小文件后续也可以直接读取结果无需下载) var outputConfig = new OutputConfig { GcsDestination = new GcsDestination { Uri = "gs://你的GCS桶名/结果输出前缀/" }, // 单个输出JSON最多包含的页数,最大支持100 BatchSize = 10 }; // 构造并提交异步处理请求 var request = new AsyncAnnotateFileRequest { InputConfig = inputConfig, Features = { feature }, OutputConfig = outputConfig }; Operation<AsyncBatchAnnotateFilesResponse, OperationMetadata> operation = await client.AsyncBatchAnnotateFilesAsync(new[] { request }); // 等待作业完成,大文件可以自行实现轮询逻辑替代阻塞等待 await operation.PollUntilCompletedAsync();
步骤2:解析检测结果提取表格内容
作业完成后,你指定的GCS输出路径下会生成output-1-to-10.json这类结构的结果文件,下载后按以下逻辑解析即可:
using Google.Cloud.Vision.V1; using Google.Protobuf; // 读取结果JSON var resultJson = File.ReadAllText(@"下载的结果JSON路径.json"); var batchResponse = BatchAnnotateFilesResponse.Parser.ParseJson(resultJson); // 遍历所有页面的标注结果 foreach (var fileResp in batchResponse.Responses) { var fullText = fileResp.FullTextAnnotation; // 筛选所有表格类型的Block var tableBlocks = fullText.Pages .SelectMany(p => p.Blocks) .Where(b => b.BlockType == Block.Types.BlockType.Table); foreach (var tableBlock in tableBlocks) { var table = tableBlock.Table; Console.WriteLine($"匹配到表格:{table.HeaderRows.Count}行表头,{table.BodyRows.Count}行表体"); // 拼接表格内容,可自行调整输出格式 foreach (var row in table.HeaderRows.Concat(table.BodyRows)) { var rowCells = row.Cells .Select(cell => // 直接取单元格文本,如需保留格式可以用fullText.Text.Substring(cell.TextOffset, cell.TextLength) string.Join("", cell.Words.SelectMany(w => w.Symbols).Select(s => s.Text)).Trim() ); Console.WriteLine(string.Join("\t", rowCells)); } } }
注意事项
- 单份PDF大小不能超过2GB,页数上限为2000页
- 异步作业执行时间和PDF页数正相关,大文件建议加超时和异常重试逻辑
- 如需提取表格的合并单元格、边框等属性,可以直接读取
Table、TableRow、TableCell对象的对应字段获取
内容的提问来源于stack exchange,提问作者Kristina Hammer
相关产品推荐
相关产品推荐

