如何在Python中实现基于视觉相似性的字符串匹配(修复Tesseract识别错误)
基于字符视觉相似性的用户名匹配方案
针对Tesseract识别时因视觉相似字符混淆导致的匹配问题,以下是两种实用的解决方案,优先采用开源工具结合自定义规则实现:
方案一:自定义视觉相似映射 + 加权编辑距离
核心是给视觉相似字符的替换设置更低的成本,让匹配算法优先识别这类错误。
1. 定义视觉易混淆字符表
先整理常见的视觉相似字符(可根据实际场景扩展):
VISUAL_SIMILAR_CHARS = { '0': ['O', 'Q'], 'O': ['0', 'Q', 'C'], 'Q': ['O', '0'], '1': ['I', 'l', 'L'], 'I': ['1', 'l', 'L'], 'l': ['1', 'I', 'L'], 'L': ['1', 'I', 'l'], '2': ['Z'], 'Z': ['2'], 'c': ['e', 'o', 'C'], 'e': ['c', 'o', 'C'], 'C': ['c', 'e', 'O'], 'p': ['q', 'b', 'P'], 'q': ['p', 'b', 'Q'], 'b': ['p', 'q', 'B'], 'P': ['p', 'q', 'B'], 'Q': ['q', 'p', 'O'], 'B': ['b', 'p', 'P'] }
2. 实现加权匹配函数
使用editdistance库(安装命令:pip install editdistance),自定义替换成本:
import editdistance def get_visual_cost(a_char, b_char): a, b = a_char.lower(), b_char.lower() if a == b: return 0 # 检查是否属于视觉相似组 if a in VISUAL_SIMILAR_CHARS and b in VISUAL_SIMILAR_CHARS[a]: return 0.2 if b in VISUAL_SIMILAR_CHARS and a in VISUAL_SIMILAR_CHARS[b]: return 0.2 return 1 def find_similar_name(target, username_list, threshold=0.5): best_match = None min_norm_score = float('inf') for name in username_list: total_cost = 0 max_len = max(len(target), len(name)) # 对齐两个字符串长度,处理长度差异 for t_char, u_char in zip(target.ljust(max_len), name.ljust(max_len)): if t_char == ' ' or u_char == ' ': total_cost += 1 # 空格差异按插入/删除算 else: total_cost += get_visual_cost(t_char, u_char) # 归一化得分,让阈值更通用 norm_score = total_cost / max_len if norm_score <= threshold and norm_score < min_norm_score: min_norm_score = norm_score best_match = name return best_match
方案二:结合拼写检查工具 + 自定义混淆规则
用pyspellchecker库,将视觉相似字符添加到自定义候选,让拼写检查优先匹配这类视觉混淆的用户名:
from spellchecker import SpellChecker import editdistance # 初始化拼写检查器,加载全量用户名作为词典 spell = SpellChecker(language=None) spell.word_frequency.load_words(username_list) # 批量添加视觉相似字符的候选映射 for char, similars in VISUAL_SIMILAR_CHARS.items(): for sim_char in similars: spell.word_frequency.add_candidate(char, sim_char) def find_similar_name(target, username_list, threshold=1): candidates = spell.candidates(target) if not candidates: return None # 筛选编辑距离符合阈值的最优匹配 sorted_candidates = sorted(candidates, key=lambda x: editdistance.eval(target, x)) for candidate in sorted_candidates: if editdistance.eval(target, candidate) <= threshold: return candidate return None
使用示例
# 模拟OCR识别结果 name_list = ["poce", "1ohn", "Zoe"] # 全量用户名库 username_list = ["pocc", "john", "zoe", "alice"] for name in name_list: match = find_similar_name(name, username_list) if match: print(f"{name} → {match}")
输出:
poce → pocc 1ohn → john Zoe → zoe
内容的提问来源于stack exchange,提问作者Sparkling Marcel
相关产品推荐
相关产品推荐

