如何使用Python从关键词列表中分离Phonetic、Word Break和Word Join类关键词
用Python分类处理关键词:拆分、音近纠错、合并
问题定义
我们需要把输入的关键词列表分成三类:
- Word Break:字母与数字连写的关键词,拆分后恢复正确格式(如
rice1kg→rice 1kg) - Phonetic:因发音相近导致的拼写错误(如
oliv oil→olive oil) - Word Join:正确单词被错误拆分的情况(如
Leg umes→legumes)
实现思路与代码
首先安装依赖库:
pip install pyenchant fuzzywuzzy python-Levenshtein phonetics
完整实现代码:
import re import enchant from fuzzywuzzy import process from phonetics import soundex # 初始化英文拼写检查器 spell_checker = enchant.Dict("en_US") def classify_keyword(keyword): keyword_clean = keyword.strip().lower() # 1. 识别Word Break:字母与数字连写场景 word_break_match = re.match(r"([a-zA-Z]+)(\d+[a-zA-Z]?)", keyword) if word_break_match: corrected = f"{word_break_match.group(1)} {word_break_match.group(2)}" return f"{corrected} - word break" # 2. 识别Word Join:拆分的两部分合并为有效单词 split_parts = keyword_clean.split() if len(split_parts) == 2: joined_word = ''.join(split_parts) if spell_checker.check(joined_word): return f"{joined_word} - word join" # 3. 识别Phonetic:发音相近的拼写错误 corrected_words = [] is_phonetic_error = False for word in split_parts: if not spell_checker.check(word): # 获取拼写检查器推荐的候选词 candidates = spell_checker.suggest(word) if candidates: # 优先匹配发音相同的候选词(用Soundex编码) target_soundex = soundex(word) phonetic_matches = [c for c in candidates if soundex(c) == target_soundex] best_match = phonetic_matches[0] if phonetic_matches else process.extractOne(word, candidates)[0] corrected_words.append(best_match) is_phonetic_error = True else: corrected_words.append(word) else: corrected_words.append(word) if is_phonetic_error: corrected = ' '.join(corrected_words) return f"{corrected} - phonetic" # 非以上三类的关键词返回原内容(如示例中的oil、cooking oil) return keyword # 测试示例输入 input_keywords = [ "rice1kg", "oil", "cooking oil", "oliv oil", "flour5kg", "buther", "baking povder", "Leg umes" ] # 输出分类结果 for kw in input_keywords: result = classify_keyword(kw) if "-" in result: print(result)
运行结果
rice 1kg - word break olive oil - phonetic flour 5kg - word break butter - phonetic baking powder - phonetic legumes - word join
逻辑说明
- Word Break检测:通过正则表达式匹配字母+数字的连写模式,拆分后添加空格。
- Word Join检测:将拆分的两个部分合并,检查是否为有效英文单词,符合则归类为合并型。
- Phonetic检测:先识别错误拼写,再通过Soundex发音编码筛选发音相近的候选词,无发音匹配时用模糊匹配取最接近的正确拼写。
内容的提问来源于stack exchange,提问作者riya23
相关产品推荐
相关产品推荐

