如何扩展OCR处理文件类型并添加Tesseract文本提取进度条
Hey there! Let's tackle your two OCR requirements one by one—adding PDF support and showing a progress bar during text extraction. Here's how to modify your existing WinForms code to make it work:
1. Add PDF File Support
Tesseract can only process image files natively, so we need to convert each page of a PDF into an image first. The PdfiumViewer library is a reliable, free option for this task in .NET.
Quick Setup:
- Install the NuGet package first:
Install-Package PdfiumViewer(or search for it in the NuGet Package Manager) - Update your code to detect PDF files and render each page as a
Pixobject (the format Tesseract uses)
2. Add a Progress Bar for Tesseract Extraction
OCR is CPU-intensive, so we'll use a BackgroundWorker to run extraction in the background (avoiding UI freezes) and report progress to update the bar.
Full Modified Code
private void btOCR_Click(object sender, EventArgs e) { if (openFileDialog.ShowDialog() == DialogResult.OK) { var filePath = openFileDialog.FileName; var pixList = new List<Pix>(); // Handle PDF vs image files if (Path.GetExtension(filePath).Equals(".pdf", StringComparison.OrdinalIgnoreCase)) { // Render each PDF page to a Pix (300 DPI for good OCR accuracy) using (var pdfDocument = PdfiumViewer.PdfDocument.Load(filePath)) { for (int pageIndex = 0; pageIndex < pdfDocument.PageCount; pageIndex++) { using (var bitmap = pdfDocument.Render(pageIndex, 300, 300, PdfiumViewer.PdfRenderFlags.Annotations)) { pixList.Add(Pix.FromBitmap(bitmap)); } } } } else { // Load image files directly pixList.Add(Pix.LoadFromFile(filePath)); } // Set up BackgroundWorker for progress tracking var bgWorker = new BackgroundWorker(); bgWorker.WorkerReportsProgress = true; bgWorker.DoWork += (s, args) => { var totalPages = pixList.Count; var extractedText = new StringBuilder(); // Initialize Tesseract engine (adjust tessdata path and language as needed) using (var ocrEngine = new TesseractEngine(@"tessdata", "eng", EngineMode.Default)) { for (int i = 0; i < totalPages; i++) { using (var page = ocrEngine.Process(pixList[i])) { extractedText.AppendLine($"--- Page {i+1} ---"); extractedText.AppendLine(page.GetText()); } // Calculate and report progress percentage int progress = (i + 1) * 100 / totalPages; bgWorker.ReportProgress(progress); } } args.Result = extractedText.ToString(); }; // Update progress bar in UI thread bgWorker.ProgressChanged += (s, args) => { progressBar1.Value = args.ProgressPercentage; }; // Handle completion or errors bgWorker.RunWorkerCompleted += (s, args) => { if (args.Error != null) { MessageBox.Show($"OCR failed: {args.Error.Message}", "Error", MessageBoxButtons.OK, MessageBoxIcon.Error); } else { // Display extracted text (replace with your actual text control) textBoxOCRResult.Text = args.Result.ToString(); } // Reset progress bar progressBar1.Value = 0; }; // Start the background task bgWorker.RunWorkerAsync(); } }
Key Tips:
- Ensure you have the Tesseract NuGet package installed:
Install-Package Tesseract - The
tessdatafolder (with language packs) should be in your app's output directory, or use an absolute path in theTesseractEngineconstructor - Add a
ProgressBarcontrol namedprogressBar1to your form, set itsMinimumto 0 andMaximumto 100 - Adjust the language code (e.g., "chi_sim" for simplified Chinese) and DPI value as needed for your use case
内容的提问来源于stack exchange,提问作者user9522816
相关产品推荐
相关产品推荐

