基于掩码与正则匹配最优编码对应集合的Python代码问题
问题:匹配最优掩码的编码集合查询
问题背景
给定包含掩码与对应集合的数组,其中X代表任意数字:
input_array = [{"code": "XXXX10", "collection": "one"}, {"code": "XXX610", "collection": "two"}, {"code": "XXXX20", "collection": "three"}]
需要编写函数接收6位编码,返回匹配最精确掩码的集合值。比如:
- 输入
010610应返回two(而非先匹配到的one) - 输入
000710应返回one
现有代码的问题
当前遍历数组时,找到第一个匹配的掩码就直接返回,导致更精确的掩码(含X更少的)被忽略。比如010610会先匹配XXXX10返回one,不符合预期。
现有代码:
import re def get_collection_from_code(analysis_code): for collection in input_array: actual_code = collection["code"] mask_to_re = actual_code.replace("X", "[\d\D]") pattern = re.compile("^" + mask_to_re + "$") if pattern.match(analysis_code): print("Found collection '" + str(collection["collection"]) + "' for code: " + str(analysis_code)) return collection["collection"] res = get_collection_from_code("010610") print(res)
解决方案
核心思路:先筛选所有匹配的掩码,再从中选出X数量最少的(最精确的)返回对应集合。
修改后的代码:
import re input_array = [{"code": "XXXX10", "collection": "one"}, {"code": "XXX610", "collection": "two"}, {"code": "XXXX20", "collection": "three"}] def get_collection_from_code(analysis_code): # 收集所有匹配的掩码项 matched_items = [] for item in input_array: mask = item["code"] # 把X替换为数字匹配(符合题目中X代表任意数字的定义) re_pattern = re.compile("^" + mask.replace("X", "\d") + "$") if re_pattern.match(analysis_code): matched_items.append(item) if not matched_items: return None # 无匹配时返回None,可按需调整 # 按掩码中X的数量排序,X越少越靠前 matched_items.sort(key=lambda x: x["code"].count("X")) best_match = matched_items[0] print(f"Found collection '{best_match['collection']}' for code: {analysis_code}") return best_match["collection"] # 测试示例 print(get_collection_from_code("010610")) # 输出 two print(get_collection_from_code("010010")) # 输出 one print(get_collection_from_code("123420")) # 输出 three
说明
- 修正正则:将
[\d\D]改为\d,严格匹配数字,符合题目中X的定义 - 先收集所有匹配项再排序,确保最精确的掩码被优先选中
- 增加无匹配场景的处理,可根据实际需求调整返回逻辑
内容的提问来源于stack exchange,提问作者Avián
相关产品推荐
相关产品推荐

