基于OpenCV的卡通字体多数字识别问题求助
解决多数字卡通字体识别的方案
我完全懂你的困扰——单数字用模板匹配准确率拉满,但一碰到多数字就直接失效,不管是做两位数模板还是限制搜索范围都没用。核心问题其实是多数字场景下不能直接用整图匹配,得先把每个数字单独处理,或者用滑动窗口遍历的方式逐个识别。下面给你几个可行的解决方案:
方案1:先分割数字,再逐个匹配(最推荐)
这是最稳妥的思路:先把多数字图片里的每个数字单独切出来,再复用你现有的单数字模板匹配逻辑识别每个数字。
步骤1:实现数字区域分割
基于你已有的二值化处理,用cv2.findContours找到每个数字的轮廓,过滤噪声后用外接矩形抠出单个数字:
def split_numbers(processed_img): # 寻找数字轮廓(只找最外层轮廓) contours, _ = cv2.findContours(processed_img, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) # 按轮廓的x坐标排序,保证识别顺序是从左到右 contours = sorted(contours, key=lambda c: cv2.boundingRect(c)[0]) number_imgs = [] for cnt in contours: x, y, w, h = cv2.boundingRect(cnt) # 过滤过小的噪声轮廓(阈值根据你的图片尺寸调整) if w > 10 and h > 10: num_img = processed_img[y:y+h, x:x+w] # 统一缩放到和模板一致的尺寸(比如模板是50x80,就把分割出的数字也缩成这个大小) num_img = cv2.resize(num_img, (50, 80)) number_imgs.append(num_img) return number_imgs
步骤2:逐个识别分割出的数字
修改你的process_num函数,先分割数字,再用模板匹配逐个识别,最后拼接结果:
def process_num(): test = cv2.imread('test.png') test = process_image(test, 11) test = cv2.resize(test, (0,0), fx=2.2, fy=2.2) # 分割出每个单独的数字 num_imgs = split_numbers(test) result = [] # 提前加载并预处理所有数字模板,避免重复IO templates = [] for i in range(10): img = cv2.imread('char-{}.png'.format(i)) img = process_lib(img, i) # 统一模板尺寸和分割后的数字一致 img = cv2.resize(img, (50, 80)) templates.append((i, img)) # 逐个匹配识别 for num_img in num_imgs: scores = [] for digit, template in templates: score = match_char(template, num_img) scores.append((score, digit)) # 取匹配得分最高的数字 best_score, best_digit = max(scores, key=lambda x: x[0]) result.append(str(best_digit)) print(f"匹配到数字{best_digit},匹配得分:{best_score:.4f}") print(f"最终识别结果:{''.join(result)}")
关键注意点
- 必须保证分割出的数字尺寸和模板尺寸完全一致,否则模板匹配会失效;
- 一定要按轮廓的x坐标排序,不然识别结果的顺序会混乱;
- 噪声过滤的阈值(
w>10, h>10)需要根据你的实际图片尺寸调整。
方案2:滑动窗口式模板匹配(适合数字间距固定的场景)
如果你的多数字图片中,数字之间的间距比较固定,也可以用滑动窗口的方式,用单数字模板在整张图上滑动,找到所有匹配度高的区域,再按位置排序得到结果:
def sliding_window_match(processed_img, template, threshold=0.8): h, w = template.shape img_h, img_w = processed_img.shape matches = [] # 遍历所有可能的窗口位置 for y in range(img_h - h + 1): for x in range(img_w - w + 1): window = processed_img[y:y+h, x:x+w] score = match_char(template, window) if score > threshold: matches.append((x, y, score)) # 去重:同一个数字可能被多个窗口匹配,保留得分最高的那个 matches = sorted(matches, key=lambda x: -x[2]) final_matches = [] used = set() for x, y, score in matches: if (x, y) not in used: final_matches.append((x, score)) # 标记周围区域为已使用,避免重复匹配 for dx in range(-5, 5): for dy in range(-5, 5): used.add((x+dx, y+dy)) # 按x坐标排序,得到从左到右的数字顺序 final_matches.sort(key=lambda x: x[0]) return final_matches
在主函数中,你可以对每个数字模板调用这个滑动窗口匹配函数,收集所有符合阈值的结果,最后合并排序得到最终的数字序列。
先修正你现有代码的小Bug
另外注意到你的match_char函数有个笔误,return max_va应该是return max_val,这个小错误会导致代码运行报错,先修正它:
def match_char(im1, im2): result = cv2.matchTemplate(im1, im2, cv2.TM_CCOEFF_NORMED) min_val, max_val, min_loc, max_loc = cv2.minMaxLoc(result) return max_val # 这里之前少了一个字母l
内容的提问来源于stack exchange,提问作者Friendier Trembley
相关产品推荐
相关产品推荐

