You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Aspose PDF与OCR提取OCR PDF表格参数设置无效果求助

如何用Aspose OCR针对性提取扫描PDF中的表格内容

我尝试使用Aspose PDF和Aspose OCR提取单页扫描PDF中的表格内容,编写了处理代码,但无论是否设置recognitionSettings.DetectAreasMode = Aspose.OCR.DetectAreasMode.Table;,提取结果完全一致,无法针对性提取表格内容,求解决。

处理代码

[HttpPost("read-pdf-aspose")]
public async Task<StandardResponse<string>> ReadPdfAspose([FromBody] PdfFilePathRequest request)
{
    try
    {
        if (string.IsNullOrWhiteSpace(request.PdfFilePath))
        {
            return new StandardResponse<string>
            {
                Status = false,
                Message = "PDF file path cannot be empty."
            };
        }

        // Create temporary directory for images
        string tempPath = Path.Combine(Path.GetTempPath(), Guid.NewGuid().ToString());
        Directory.CreateDirectory(tempPath);

        try
        {
            // Load PDF document
            using (Document pdfDocument = new Document(request.PdfFilePath))
            {
                // Initialize OCR engine
                Aspose.OCR.AsposeOcr recognitionEngine = new Aspose.OCR.AsposeOcr();
                StringBuilder extractedText = new StringBuilder();

                // Process each page
                for (int pageIndex = 1; pageIndex <= pdfDocument.Pages.Count; pageIndex++)
                {
                    var page = pdfDocument.Pages[pageIndex];

                    // Save images from the page
                    for (int imgIndex = 1; imgIndex <= page.Resources.Images.Count; imgIndex++)
                    {
                        string imagePath = Path.Combine(tempPath, $"page_{pageIndex}_img_{imgIndex}.png");

                        // Extract and save the image
                        using (FileStream imageStream = new(imagePath, FileMode.Create))
                        {
                            page.Resources.Images[imgIndex].Save(imageStream);
                        }

                        // Set up OCR for table detection
                        Aspose.OCR.OcrInput input = new Aspose.OCR.OcrInput(Aspose.OCR.InputType.SingleImage);
                        input.Add(imagePath);

                        // Configure to detect tables
                        Aspose.OCR.RecognitionSettings recognitionSettings = new Aspose.OCR.RecognitionSettings();
                         recognitionSettings.DetectAreasMode = Aspose.OCR.DetectAreasMode.Table;

                        // Perform recognition
                        Aspose.OCR.OcrOutput results = recognitionEngine.Recognize(input, recognitionSettings);

                        // Collect recognized text
                        extractedText.AppendLine($"--- Table content from Page {pageIndex}, Image {imgIndex} ---");
                        foreach (Aspose.OCR.RecognitionResult result in results)
                        {
                            extractedText.AppendLine(result.RecognitionText);
                        }
                        extractedText.AppendLine();
                    }
                }

                return new StandardResponse<string>
                {
                    Status = true,
                    Message = "Tables extracted successfully from PDF using Aspose",
                    Data = extractedText.ToString()
                };
            }
        }
        finally
        {
            // Cleanup: Delete temporary directory and files
            if (Directory.Exists(tempPath))
            {
                Directory.Delete(tempPath, true);
            }
        }
    }
    catch (Exception ex)
    {
        _logger.LogError(ex, "Error processing PDF with Aspose");
        return new StandardResponse<string>
        {
            Status = false,
            Message = ex.Message
        };
    }
}

当前输出结果

无论是否启用DetectAreasMode.Table,返回结果均为:

{
    "status": true,
    "message": "Tables extracted successfully from PDF using Aspose",
    "data": "--- Table content from Page 1, Image 1 ---
The Main Table reguired for evaluation
Header1 Second Header Third Header Fourth Header
First First Sample First Second First Third Sample First Fourth
Sample Sample
Second First Second Second Second Third Second Fourth
Sample Sample Sample Sample
Third First Sample Third Second Third Third Sample Third Fourth
Sample Sample
Sample text for differentiation between Table data and non table data
Ipsum has been the industry's standard dummy text ever since the 1500s,when an
unknown printer took a galley of type and scrambled it to make a type specimen book. It
has survived not only five centuries,but also the leap into electronic typesetting,remaining
essentiallyl

"
}

解决方案

1. 改用结构化表格提取逻辑

你当前只读取了RecognitionResult.RecognitionText,这是整页的纯文本,不会区分表格和普通文本。Aspose.OCR专门提供了RecognitionResult.Tables属性来获取结构化的表格数据,修改结果收集部分的代码:

// 替换原来的结果收集代码
extractedText.AppendLine($"--- 表格内容 页{pageIndex}, 图{imgIndex} ---");
foreach (Aspose.OCR.RecognitionResult result in results)
{
    // 遍历识别到的每个表格
    foreach(var table in result.Tables)
    {
        // 逐行输出表格单元格,用制表符分隔方便后续解析
        foreach(var row in table.Rows)
        {
            extractedText.AppendLine(string.Join("\t", row.Cells));
        }
        extractedText.AppendLine();
    }
}

2. 修正PDF页面转图片逻辑

如果是扫描PDF,整个页面是一张图片,你当前提取资源图片的逻辑可能没问题,但如果是包含嵌入图片+文本的混合PDF,建议直接把整个页面转成图片,确保不遗漏内容:

// 替换原来的提取图片循环,将整个PDF页面转为图片
var resolution = new Resolution(300); // 300DPI保证识别精度
var pngDevice = new PngDevice(resolution);
string imagePath = Path.Combine(tempPath, $"page_{pageIndex}.png");
pngDevice.Process(page, imagePath);

3. 优化识别参数

除了设置DetectAreasMode.Table,补充以下参数提升表格识别效果:

Aspose.OCR.RecognitionSettings recognitionSettings = new Aspose.OCR.RecognitionSettings();
recognitionSettings.DetectAreasMode = Aspose.OCR.DetectAreasMode.Table;
recognitionSettings.AutoSkew = true; // 自动校正倾斜
recognitionSettings.ContrastCorrection = true; // 自动增强对比度
recognitionSettings.IgnoreNoise = true; // 忽略背景噪点

内容的提问来源于stack exchange,提问作者Kaif Khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 02:53:12