如何用Python生成OCR误识别场景下所有字符替换的排列组合
实现思路
核心逻辑为笛卡尔积遍历:
- 对输入字符串的每个字符,生成该位置所有可能的取值:若字符存在于误分类映射规则中,取值包含
原字符 + 所有误识别结果;若不存在,取值仅为原字符本身 - 对所有位置的取值列表做笛卡尔积运算,将每组运算结果拼接为字符串,即为所有可能的误分类替换结果
Python实现代码
import itertools from typing import List, Dict def generate_ocr_permutations(input_str: str, misclassification_rule: Dict[str, List[str]]) -> List[str]: # 构造每个位置的候选值列表 char_candidates = [] for char in input_str: if char in misclassification_rule: # 候选包含原字符 + 所有误识别结果 candidates = [char] + misclassification_rule[char] else: candidates = [char] char_candidates.append(candidates) # 求笛卡尔积并拼接为字符串 permutations = [] for group in itertools.product(*char_candidates): permutations.append(''.join(group)) # 可选:如果存在不同替换路径得到相同结果的场景,可打开下一行去重 # permutations = list(set(permutations)) return permutations
测试示例
# 代入你给出的规则和输入测试 rule = {'1':['J'], '2':['Z','J']} input_str = "AB1CD2" result = generate_ocr_permutations(input_str, rule) print(result) # 输出结果:['AB1CD2', 'AB1CDZ', 'AB1CDJ', 'ABJCD2', 'ABJCDZ', 'ABJCDJ'] 共6条,和预期完全一致
补充说明
- 如果需要处理大小写不敏感的场景,可以提前将输入字符串和映射规则的键统一转为大写/小写再处理
- 当规则比较复杂时,结果数量会呈指数级增长,可根据业务需要增加长度阈值限制
内容的提问来源于stack exchange,提问作者Siddharth Shivkumar
相关产品推荐
相关产品推荐

