You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中实现基于视觉相似性的字符串匹配(修复Tesseract识别错误)

基于字符视觉相似性的用户名匹配方案

针对Tesseract识别时因视觉相似字符混淆导致的匹配问题,以下是两种实用的解决方案,优先采用开源工具结合自定义规则实现:

方案一:自定义视觉相似映射 + 加权编辑距离

核心是给视觉相似字符的替换设置更低的成本,让匹配算法优先识别这类错误。

1. 定义视觉易混淆字符表

先整理常见的视觉相似字符(可根据实际场景扩展):

VISUAL_SIMILAR_CHARS = {
    '0': ['O', 'Q'],
    'O': ['0', 'Q', 'C'],
    'Q': ['O', '0'],
    '1': ['I', 'l', 'L'],
    'I': ['1', 'l', 'L'],
    'l': ['1', 'I', 'L'],
    'L': ['1', 'I', 'l'],
    '2': ['Z'],
    'Z': ['2'],
    'c': ['e', 'o', 'C'],
    'e': ['c', 'o', 'C'],
    'C': ['c', 'e', 'O'],
    'p': ['q', 'b', 'P'],
    'q': ['p', 'b', 'Q'],
    'b': ['p', 'q', 'B'],
    'P': ['p', 'q', 'B'],
    'Q': ['q', 'p', 'O'],
    'B': ['b', 'p', 'P']
}

2. 实现加权匹配函数

使用editdistance库(安装命令:pip install editdistance),自定义替换成本:

import editdistance

def get_visual_cost(a_char, b_char):
    a, b = a_char.lower(), b_char.lower()
    if a == b:
        return 0
    # 检查是否属于视觉相似组
    if a in VISUAL_SIMILAR_CHARS and b in VISUAL_SIMILAR_CHARS[a]:
        return 0.2
    if b in VISUAL_SIMILAR_CHARS and a in VISUAL_SIMILAR_CHARS[b]:
        return 0.2
    return 1

def find_similar_name(target, username_list, threshold=0.5):
    best_match = None
    min_norm_score = float('inf')
    for name in username_list:
        total_cost = 0
        max_len = max(len(target), len(name))
        # 对齐两个字符串长度,处理长度差异
        for t_char, u_char in zip(target.ljust(max_len), name.ljust(max_len)):
            if t_char == ' ' or u_char == ' ':
                total_cost += 1  # 空格差异按插入/删除算
            else:
                total_cost += get_visual_cost(t_char, u_char)
        # 归一化得分,让阈值更通用
        norm_score = total_cost / max_len
        if norm_score <= threshold and norm_score < min_norm_score:
            min_norm_score = norm_score
            best_match = name
    return best_match

方案二:结合拼写检查工具 + 自定义混淆规则

用pyspellchecker库,将视觉相似字符添加到自定义候选,让拼写检查优先匹配这类视觉混淆的用户名:

from spellchecker import SpellChecker
import editdistance

# 初始化拼写检查器,加载全量用户名作为词典
spell = SpellChecker(language=None)
spell.word_frequency.load_words(username_list)

# 批量添加视觉相似字符的候选映射
for char, similars in VISUAL_SIMILAR_CHARS.items():
    for sim_char in similars:
        spell.word_frequency.add_candidate(char, sim_char)

def find_similar_name(target, username_list, threshold=1):
    candidates = spell.candidates(target)
    if not candidates:
        return None
    # 筛选编辑距离符合阈值的最优匹配
    sorted_candidates = sorted(candidates, key=lambda x: editdistance.eval(target, x))
    for candidate in sorted_candidates:
        if editdistance.eval(target, candidate) <= threshold:
            return candidate
    return None

使用示例

# 模拟OCR识别结果
name_list = ["poce", "1ohn", "Zoe"]
# 全量用户名库
username_list = ["pocc", "john", "zoe", "alice"]

for name in name_list:
    match = find_similar_name(name, username_list)
    if match:
        print(f"{name} → {match}")

输出:

poce → pocc
1ohn → john
Zoe → zoe

内容的提问来源于stack exchange,提问作者Sparkling Marcel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 21:07:33