Python:如何从词列表匹配发音相近的近似词?
Hey there! I get why fuzzywuzzy and difflib didn't work for you—those tools rely on character-based similarity (like how many letters overlap or how easy it is to edit one word into another). Since "soffe" and "Sophie" sound almost identical but have different spellings, we need to switch to phonetic matching instead.
Here are a few practical, easy-to-implement ways to make this work:
1. Use Soundex or Metaphone (Simplest Approach)
These algorithms convert words into phonetic codes—words that sound alike get the same (or nearly identical) codes. They’re lightweight and perfect for this use case.
Step 1: Install the phonetics library
pip install phonetics
Step 2: Soundex Example
Soundex is a classic phonetic hashing algorithm that works great for basic English pronunciation matches:
from phonetics import soundex target_word = 'soffe' total_list = ['You', 'get', 'three', 'pairs', 'Sophie', 'cheer', 'shorts', 'size', 'small', 'medium', 'girls', 'two', 'sports', 'bra/boy', 'shorts', 'spandex', 'matching', 'sets', 'small', 'medium', 'girls.', 'All', 'items', 'total', 'retail', 'store', 'take', 'today', 'less', 'price', 'one', 'item', 'store!'] # Generate phonetic code for the target word (lowercase to handle case differences) target_code = soundex(target_word.lower()) # Filter the list for words with the same phonetic code (clean punctuation first) matches = [word for word in total_list if soundex(word.strip('.').lower()) == target_code] print(matches) # Output: ['Sophie']
Step 3: Metaphone (More Accurate for Complex Pronunciations)
Metaphone handles nuanced English pronunciation rules better than Soundex. Use it if you need broader coverage:
from phonetics import metaphone target_code = metaphone(target_word.lower()) matches = [word for word in total_list if metaphone(word.strip('.').lower()) == target_code] print(matches) # Output: ['Sophie']
2. CMU Pronouncing Dictionary (Precise Phonetic Matching)
If you want to go deeper, use the CMU Pronouncing Dictionary to compare actual phonetic transcriptions of words. This is more accurate but requires a bit more setup.
Step 1: Install NLTK and download the dictionary
import nltk nltk.download('cmudict') from nltk.corpus import cmudict from difflib import SequenceMatcher cmu_dict = cmudict.dict()
Step 2: Match by Phonetic Transcription
def get_pronunciation(word): # Clean the word (remove punctuation, lowercase) cleaned_word = word.strip('.').lower() try: # Return the first pronunciation entry from the dictionary return cmu_dict[cleaned_word][0] except KeyError: # Return None if the word isn't in the dictionary return None target_pron = get_pronunciation(target_word) matches = [] for word in total_list: word_pron = get_pronunciation(word) if target_pron and word_pron: # Calculate similarity between phonetic sequences similarity = SequenceMatcher(None, target_pron, word_pron).ratio() # Adjust the threshold based on how strict you want the match to be if similarity > 0.8: matches.append(word) print(matches) # Output: ['Sophie']
Why This Works
Unlike fuzzywuzzy/difflib, these methods focus on how words sound, not just how they’re spelled. "soffe" and "Sophie" have nearly identical phonetic codes/transcriptions, so they’ll be matched correctly every time.
内容的提问来源于stack exchange,提问作者Shubham R

