You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C#中实现渔民支票图片16位波斯数字渔民ID精准定位与提取的问题求助

C#中实现渔民支票图片16位波斯数字渔民ID精准定位与提取的问题求助

大家好,我现在需要实现一个功能:处理渔民支票的图片,首先要精准检测出图片中渔民ID的位置,然后提取出这串包含16位波斯数字的ID并输出。我自己写了一段基于Tesseract的C#代码,但目前遇到的问题是OCR没办法完整提取所有数字,想请各位开发者帮忙分析下问题所在,或者给出优化建议?

以下是我目前的代码:

using System;
using System.Drawing;
using System.Text.RegularExpressions;
using Tesseract;
using Rectangle = System.Drawing.Rectangle;
using System.IO; // 补充原代码缺失的命名空间,用于文件写入操作

// Load the image
Bitmap image = new Bitmap("C:\\Users\\4020302\\Iron\\f.PNG");

var pix2 = Pix.LoadFromFile("C:\\Users\\4020302\\Iron\\f.PNG");

var adadha = "";

using (var engine = new TesseractEngine(@"./tessdata", "fas", EngineMode.Default))
{
    // Create an instance of Page
    using (var page = engine.Process(pix2))
    {
        // Get an iterator over the text elements
        using (var iter = page.GetIterator())
        {
            // Loop through the iterator
            do
            {
                // Get the text of the current element
                var text = iter.GetText(PageIteratorLevel.Word);

                // Check if the text matches the sayad pattern
                var regex = new Regex("^[\u06F0-\u06F90-9]+$");
                var match = regex.Match(text);

                if (match.Success)
                {
                    // Get the bounding box of the current element
                    Rect res;
                    var rect = iter.TryGetBoundingBox(PageIteratorLevel.Word, out res);

                    // Crop the image using the bounding box
                    var subImage = CropImage(image, res);

                    // Save or display the sub-image
                    subImage.Save("sayad_id.png");

                    Console.WriteLine("The sayad id in the image is: " + text);
                    Console.WriteLine("The location of the sayad id is: " + rect);

                    adadha += text;
                }
            } while (iter.Next(PageIteratorLevel.Word));
        }
    }
}

File.WriteAllText("C:\\Users\\4020302\\Iron\\f.txt", adadha);

// A method to crop an image using a rectangle
static Bitmap CropImage(Bitmap source, Rect area)
{
    Bitmap target = new Bitmap(area.Width, area.Height);
    using (Graphics g = Graphics.FromImage(target))
    {
        g.DrawImage(source, new Rectangle(0, 0, target.Width, target.Height),
            area.X1, area.Y1, area.Width, area.Height,
            GraphicsUnit.Pixel);
    }
    return target;
}

目前的问题是,这段代码只能提取出部分波斯数字,没办法完整获取16位的ID内容。我已经确保tessdata目录下有波斯语的训练数据,但还是没解决问题,想问问大家有没有什么改进的思路?比如图片预处理(灰度化、二值化、降噪)、Tesseract的参数调整,或者更精准的定位逻辑?

备注:内容来源于stack exchange,提问作者Aidin Barmalaei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.21 15:08:08