求Python中Reverse Soundex算法的实现方案或现成函数
Reverse Soundex 的 Python 实现
Reverse Soundex 是标准 Soundex 的变体,核心差异是从单词末尾开始向前处理字符,编码规则与标准 Soundex 兼容。目前没有广泛维护的现成 Python 库提供该算法,以下是手动实现方案:
实现代码
def reverse_soundex(word): # 处理空输入或非字符串输入 if not isinstance(word, str) or len(word) == 0: return "" # 统一转为大写,兼容大小写输入 word = word.upper() # 沿用标准Soundex的辅音数字映射规则 soundex_map = { 'B': '1', 'F': '1', 'P': '1', 'V': '1', 'C': '2', 'G': '2', 'J': '2', 'K': '2', 'Q': '2', 'S': '2', 'X': '2', 'Z': '2', 'D': '3', 'T': '3', 'L': '4', 'M': '5', 'N': '5', 'R': '6' } # 保留单词最后一个字符作为编码首字符 last_char = word[-1] code = [last_char] # 从倒数第二个字符向前遍历处理 for char in reversed(word[:-1]): # 忽略元音、半元音及h/w if char in {'A', 'E', 'I', 'O', 'U', 'Y', 'H', 'W'}: continue mapped = soundex_map.get(char) if not mapped: continue # 跳过连续重复的映射数字 if mapped != code[-1]: code.append(mapped) # 调整编码长度为4位,不足补0,过长截断 while len(code) < 4: code.append('0') return ''.join(code[:4])
代码说明
- 输入兼容:自动转为大写,处理空输入、非字符串等边界情况
- 规则匹配:完全复用标准 Soundex 的辅音映射逻辑,仅遍历方向相反
- 去重优化:避免连续重复的数字编码,保证结果简洁
- 长度标准化:最终编码固定为4位,与标准 Soundex 格式统一
示例验证
print(reverse_soundex("Smith")) # 输出 H530 print(reverse_soundex("Smyth")) # 输出 H530(语音匹配特性下与Smith编码一致) print(reverse_soundex("Johnson")) # 输出 N250
内容的提问来源于stack exchange,提问作者Shubham Tomar
相关产品推荐
相关产品推荐

