Fuzzywuzzy process.extractBests结果异常,是否需预处理优化?
Yep, preprocessing your text and tweaking the matching parameters will absolutely fix this issue! Let's walk through why your current results are off, and how to get those 4 "St Paul" schools to show up at the top.
Why Your Current Results Are Unexpected
The main problems here are:
- Punctuation mismatch: Your search query uses
"St. Paul"(with a dot), but all the target schools use"St Paul"(no dot). This tiny difference lowers the fuzzy match score for the schools you actually want. - Default algorithm behavior: The default
WRatioscorer looks at full-string similarity. The schools that are showing up (like"St Clare's Girl's School") are shorter and closer in length to"St. Paul", so they get higher scores than the longer "St Paul" schools with extra suffixes (like"St Paul's Co-educational College").
Step 1: Standardize Text with Preprocessing
First, we'll clean up both the search query and school list to eliminate punctuation inconsistencies:
from fuzzywuzzy import process, fuzz def preprocess_text(text): # Remove dots and trim any extra whitespace return text.replace(".", "").strip() # Your original school list schList = ["Diocesan Boy's School", "Diocesan Girl's School", 'Heep Yunn School', 'La Salle College', 'Maryknoll Convent School', 'Marymount Secondary School', 'Methodist College', 'Sacred Heart Canossian College', "St Clare's Girl's School", 'St Francis Canossian College', "St Joseph's College", "St Mark's School", "St Mary's Canossian College", "St Paul's Co-educational College", "St Paul's College", "St Paul's Convent School", "St Paul's Secondary School", "St Stephen's Girl's College", 'Wah Yan College, Hong Kong', 'Wah Yan College, Kowloon', 'Ying Wa College', "Ying Wa Girl's School"] # Preprocess the school list and query cleaned_schools = [preprocess_text(school) for school in schList] cleaned_query = preprocess_text("St. Paul")
Step 2: Use a Scorer That Fits Your Needs
If you want to prioritize schools that start with (or contain) "St Paul", fuzz.partial_ratio is perfect—it focuses on partial matches instead of full-string similarity, so the suffixes on the "St Paul" schools won't drag down their scores.
Here's how to implement it:
# Get matching results with partial ratio scoring matches = process.extractBests(cleaned_query, cleaned_schools, scorer=fuzz.partial_ratio, score_cutoff=85) # Map back to the original school names (since we cleaned them) final_results = [(schList[cleaned_schools.index(cleaned_school)], score) for cleaned_school, score in matches] print(final_results)
This should return exactly the 4 schools you're expecting, with high scores (usually 90+).
If you prefer full-string matching, even just preprocessing will boost the scores of the "St Paul" schools enough to push them to the top:
matches = process.extractBests(cleaned_query, cleaned_schools, score_cutoff=80) final_results = [(schList[cleaned_schools.index(cleaned_school)], score) for cleaned_school, score in matches] print(final_results)
Bonus: Tweak Limits and Cutoffs
- Use the
limitparameter to control how many results you get (e.g.,limit=4to only return the top 4 matches). - Adjust
score_cutoffto filter out low-quality matches entirely.
内容的提问来源于stack exchange,提问作者JOSEPH CHONG

