You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于OpenCV的卡通字体多数字识别问题求助

解决多数字卡通字体识别的方案

我完全懂你的困扰——单数字用模板匹配准确率拉满,但一碰到多数字就直接失效,不管是做两位数模板还是限制搜索范围都没用。核心问题其实是多数字场景下不能直接用整图匹配,得先把每个数字单独处理,或者用滑动窗口遍历的方式逐个识别。下面给你几个可行的解决方案:

方案1:先分割数字,再逐个匹配(最推荐)

这是最稳妥的思路:先把多数字图片里的每个数字单独切出来,再复用你现有的单数字模板匹配逻辑识别每个数字。

步骤1:实现数字区域分割

基于你已有的二值化处理,用cv2.findContours找到每个数字的轮廓,过滤噪声后用外接矩形抠出单个数字:

def split_numbers(processed_img):
    # 寻找数字轮廓(只找最外层轮廓)
    contours, _ = cv2.findContours(processed_img, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
    # 按轮廓的x坐标排序,保证识别顺序是从左到右
    contours = sorted(contours, key=lambda c: cv2.boundingRect(c)[0])
    number_imgs = []
    
    for cnt in contours:
        x, y, w, h = cv2.boundingRect(cnt)
        # 过滤过小的噪声轮廓(阈值根据你的图片尺寸调整)
        if w > 10 and h > 10:
            num_img = processed_img[y:y+h, x:x+w]
            # 统一缩放到和模板一致的尺寸(比如模板是50x80,就把分割出的数字也缩成这个大小)
            num_img = cv2.resize(num_img, (50, 80))
            number_imgs.append(num_img)
    return number_imgs

步骤2:逐个识别分割出的数字

修改你的process_num函数,先分割数字,再用模板匹配逐个识别,最后拼接结果:

def process_num():
    test = cv2.imread('test.png')
    test = process_image(test, 11)
    test = cv2.resize(test, (0,0), fx=2.2, fy=2.2)
    
    # 分割出每个单独的数字
    num_imgs = split_numbers(test)
    result = []
    
    # 提前加载并预处理所有数字模板,避免重复IO
    templates = []
    for i in range(10):
        img = cv2.imread('char-{}.png'.format(i))
        img = process_lib(img, i)
        # 统一模板尺寸和分割后的数字一致
        img = cv2.resize(img, (50, 80))
        templates.append((i, img))
    
    # 逐个匹配识别
    for num_img in num_imgs:
        scores = []
        for digit, template in templates:
            score = match_char(template, num_img)
            scores.append((score, digit))
        # 取匹配得分最高的数字
        best_score, best_digit = max(scores, key=lambda x: x[0])
        result.append(str(best_digit))
        print(f"匹配到数字{best_digit},匹配得分:{best_score:.4f}")
    
    print(f"最终识别结果:{''.join(result)}")

关键注意点

  • 必须保证分割出的数字尺寸和模板尺寸完全一致,否则模板匹配会失效;
  • 一定要按轮廓的x坐标排序,不然识别结果的顺序会混乱;
  • 噪声过滤的阈值(w>10, h>10)需要根据你的实际图片尺寸调整。

方案2:滑动窗口式模板匹配(适合数字间距固定的场景)

如果你的多数字图片中,数字之间的间距比较固定,也可以用滑动窗口的方式,用单数字模板在整张图上滑动,找到所有匹配度高的区域,再按位置排序得到结果:

def sliding_window_match(processed_img, template, threshold=0.8):
    h, w = template.shape
    img_h, img_w = processed_img.shape
    matches = []
    
    # 遍历所有可能的窗口位置
    for y in range(img_h - h + 1):
        for x in range(img_w - w + 1):
            window = processed_img[y:y+h, x:x+w]
            score = match_char(template, window)
            if score > threshold:
                matches.append((x, y, score))
    
    # 去重:同一个数字可能被多个窗口匹配,保留得分最高的那个
    matches = sorted(matches, key=lambda x: -x[2])
    final_matches = []
    used = set()
    for x, y, score in matches:
        if (x, y) not in used:
            final_matches.append((x, score))
            # 标记周围区域为已使用,避免重复匹配
            for dx in range(-5, 5):
                for dy in range(-5, 5):
                    used.add((x+dx, y+dy))
    
    # 按x坐标排序,得到从左到右的数字顺序
    final_matches.sort(key=lambda x: x[0])
    return final_matches

在主函数中,你可以对每个数字模板调用这个滑动窗口匹配函数,收集所有符合阈值的结果,最后合并排序得到最终的数字序列。

先修正你现有代码的小Bug

另外注意到你的match_char函数有个笔误,return max_va应该是return max_val,这个小错误会导致代码运行报错,先修正它:

def match_char(im1, im2):
    result = cv2.matchTemplate(im1, im2, cv2.TM_CCOEFF_NORMED)
    min_val, max_val, min_loc, max_loc = cv2.minMaxLoc(result)
    return max_val  # 这里之前少了一个字母l

内容的提问来源于stack exchange,提问作者Friendier Trembley

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:32:49