Aspose PDF与OCR提取OCR PDF表格参数设置无效果求助
如何用Aspose OCR针对性提取扫描PDF中的表格内容
我尝试使用Aspose PDF和Aspose OCR提取单页扫描PDF中的表格内容,编写了处理代码,但无论是否设置recognitionSettings.DetectAreasMode = Aspose.OCR.DetectAreasMode.Table;,提取结果完全一致,无法针对性提取表格内容,求解决。
处理代码
[HttpPost("read-pdf-aspose")] public async Task<StandardResponse<string>> ReadPdfAspose([FromBody] PdfFilePathRequest request) { try { if (string.IsNullOrWhiteSpace(request.PdfFilePath)) { return new StandardResponse<string> { Status = false, Message = "PDF file path cannot be empty." }; } // Create temporary directory for images string tempPath = Path.Combine(Path.GetTempPath(), Guid.NewGuid().ToString()); Directory.CreateDirectory(tempPath); try { // Load PDF document using (Document pdfDocument = new Document(request.PdfFilePath)) { // Initialize OCR engine Aspose.OCR.AsposeOcr recognitionEngine = new Aspose.OCR.AsposeOcr(); StringBuilder extractedText = new StringBuilder(); // Process each page for (int pageIndex = 1; pageIndex <= pdfDocument.Pages.Count; pageIndex++) { var page = pdfDocument.Pages[pageIndex]; // Save images from the page for (int imgIndex = 1; imgIndex <= page.Resources.Images.Count; imgIndex++) { string imagePath = Path.Combine(tempPath, $"page_{pageIndex}_img_{imgIndex}.png"); // Extract and save the image using (FileStream imageStream = new(imagePath, FileMode.Create)) { page.Resources.Images[imgIndex].Save(imageStream); } // Set up OCR for table detection Aspose.OCR.OcrInput input = new Aspose.OCR.OcrInput(Aspose.OCR.InputType.SingleImage); input.Add(imagePath); // Configure to detect tables Aspose.OCR.RecognitionSettings recognitionSettings = new Aspose.OCR.RecognitionSettings(); recognitionSettings.DetectAreasMode = Aspose.OCR.DetectAreasMode.Table; // Perform recognition Aspose.OCR.OcrOutput results = recognitionEngine.Recognize(input, recognitionSettings); // Collect recognized text extractedText.AppendLine($"--- Table content from Page {pageIndex}, Image {imgIndex} ---"); foreach (Aspose.OCR.RecognitionResult result in results) { extractedText.AppendLine(result.RecognitionText); } extractedText.AppendLine(); } } return new StandardResponse<string> { Status = true, Message = "Tables extracted successfully from PDF using Aspose", Data = extractedText.ToString() }; } } finally { // Cleanup: Delete temporary directory and files if (Directory.Exists(tempPath)) { Directory.Delete(tempPath, true); } } } catch (Exception ex) { _logger.LogError(ex, "Error processing PDF with Aspose"); return new StandardResponse<string> { Status = false, Message = ex.Message }; } }
当前输出结果
无论是否启用DetectAreasMode.Table,返回结果均为:
{ "status": true, "message": "Tables extracted successfully from PDF using Aspose", "data": "--- Table content from Page 1, Image 1 --- The Main Table reguired for evaluation Header1 Second Header Third Header Fourth Header First First Sample First Second First Third Sample First Fourth Sample Sample Second First Second Second Second Third Second Fourth Sample Sample Sample Sample Third First Sample Third Second Third Third Sample Third Fourth Sample Sample Sample text for differentiation between Table data and non table data Ipsum has been the industry's standard dummy text ever since the 1500s,when an unknown printer took a galley of type and scrambled it to make a type specimen book. It has survived not only five centuries,but also the leap into electronic typesetting,remaining essentiallyl " }
解决方案
1. 改用结构化表格提取逻辑
你当前只读取了RecognitionResult.RecognitionText,这是整页的纯文本,不会区分表格和普通文本。Aspose.OCR专门提供了RecognitionResult.Tables属性来获取结构化的表格数据,修改结果收集部分的代码:
// 替换原来的结果收集代码 extractedText.AppendLine($"--- 表格内容 页{pageIndex}, 图{imgIndex} ---"); foreach (Aspose.OCR.RecognitionResult result in results) { // 遍历识别到的每个表格 foreach(var table in result.Tables) { // 逐行输出表格单元格,用制表符分隔方便后续解析 foreach(var row in table.Rows) { extractedText.AppendLine(string.Join("\t", row.Cells)); } extractedText.AppendLine(); } }
2. 修正PDF页面转图片逻辑
如果是扫描PDF,整个页面是一张图片,你当前提取资源图片的逻辑可能没问题,但如果是包含嵌入图片+文本的混合PDF,建议直接把整个页面转成图片,确保不遗漏内容:
// 替换原来的提取图片循环,将整个PDF页面转为图片 var resolution = new Resolution(300); // 300DPI保证识别精度 var pngDevice = new PngDevice(resolution); string imagePath = Path.Combine(tempPath, $"page_{pageIndex}.png"); pngDevice.Process(page, imagePath);
3. 优化识别参数
除了设置DetectAreasMode.Table,补充以下参数提升表格识别效果:
Aspose.OCR.RecognitionSettings recognitionSettings = new Aspose.OCR.RecognitionSettings(); recognitionSettings.DetectAreasMode = Aspose.OCR.DetectAreasMode.Table; recognitionSettings.AutoSkew = true; // 自动校正倾斜 recognitionSettings.ContrastCorrection = true; // 自动增强对比度 recognitionSettings.IgnoreNoise = true; // 忽略背景噪点
内容的提问来源于stack exchange,提问作者Kaif Khan
相关产品推荐
相关产品推荐

