求免费.NET(MVC)扫描+OCR提取指定字段存入数据库的类库
Hey there! Let's break down what you need—scanning documents from a scanner, running OCR to pull text, and extracting specific fields (like name and insurance ID) to save to a database—using free .NET MVC-compatible tools. Here's a practical, developer-focused breakdown:
If you're building a desktop .NET MVC app, the built-in Windows Image Acquisition (WIA) API is a solid free option. It lets you interact directly with scanners without extra paid libraries. You'll just need to add a reference to the WIA interop assembly (usually Interop.WIA.dll) or leverage scanning utilities in the System.Windows.Forms namespace (Windows-only, but free).
For ASP.NET MVC (web apps): Browsers can't access scanners directly, so you'll need users to scan documents locally first (via a desktop tool or free client-side JS scanning tools) then upload the scanned image to your MVC backend for processing.
Tesseract is a powerful open-source OCR engine, and its .NET wrapper Tesseract.NET is perfect for this use case. It's free, NuGet-installable, and works with both desktop and web .NET projects.
To get started:
- Install the
TesseractNuGet package in your MVC project. - Download the Tesseract language data pack (English is default) and place it in a
tessdatafolder in your project.
Once you have the raw OCR text, regex is your best bet for pulling specific fields—assuming your insurance documents have consistent formatting (e.g., Name: John Doe or Insurance ID: INS-12345).
For example, if your document has lines like:
Full Name: Jane Smith
Insurance Policy ID: INS-98765
You can craft regex patterns to match these lines and extract the values cleanly.
Here's a simplified snippet that ties scanning, OCR, field extraction, and database saving together (using EF Core for database operations):
// Step 1: Scan document (WIA example for desktop) var scanner = new WIA.DeviceManager(); var scannerDevice = scanner.DeviceInfos.Cast<WIA.DeviceInfo>() .First(d => d.Type == WIA.WiaDeviceType.ScannerDeviceType) .Connect(); var scanItem = scannerDevice.Items.Cast<WIA.Item>().First(); var imageFile = (WIA.ImageFile)scanItem.Transfer(WIA.FormatID.wiaFormatPNG); imageFile.SaveFile(@"C:\ScannedDocs\insurance_scan.png"); // Step 2: Run OCR with Tesseract using var engine = new TesseractEngine(@"./tessdata", "eng", EngineMode.Default); using var img = Pix.LoadFromFile(@"C:\ScannedDocs\insurance_scan.png"); using var page = engine.Process(img); var ocrText = page.GetText(); // Step 3: Extract Name and Insurance ID with regex var nameMatch = Regex.Match(ocrText, @"Full Name:\s*(.*)", RegexOptions.IgnoreCase); var insuranceIdMatch = Regex.Match(ocrText, @"Insurance Policy ID:\s*(.*)", RegexOptions.IgnoreCase); string extractedName = nameMatch.Success ? nameMatch.Groups[1].Value.Trim() : string.Empty; string extractedId = insuranceIdMatch.Success ? insuranceIdMatch.Groups[1].Value.Trim() : string.Empty; // Step 4: Save to database (EF Core example) using var dbContext = new YourInsuranceDbContext(); dbContext.InsuranceRecords.Add(new InsuranceRecord { FullName = extractedName, PolicyId = extractedId, ScannedDate = DateTime.Now }); await dbContext.SaveChangesAsync();
- WIA is Windows-only—if you need cross-platform support, free scanning options are limited, but you can pair cross-platform image processing libraries (like
ImageSharp) with user-uploaded scans. - Tesseract's accuracy improves with high-resolution, clean scans. Preprocess images (resize, adjust contrast) using free .NET libraries to boost OCR results.
- Refine your regex patterns with sample scans—document formatting can vary, so test with real-world examples to ensure reliable extraction.
内容的提问来源于stack exchange,提问作者lorraine chen

