You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Fuzzywuzzy process.extractBests结果异常,是否需预处理优化?

Fixing Fuzzy Matching for "St. Paul" School Results

Yep, preprocessing your text and tweaking the matching parameters will absolutely fix this issue! Let's walk through why your current results are off, and how to get those 4 "St Paul" schools to show up at the top.

Why Your Current Results Are Unexpected

The main problems here are:

  • Punctuation mismatch: Your search query uses "St. Paul" (with a dot), but all the target schools use "St Paul" (no dot). This tiny difference lowers the fuzzy match score for the schools you actually want.
  • Default algorithm behavior: The default WRatio scorer looks at full-string similarity. The schools that are showing up (like "St Clare's Girl's School") are shorter and closer in length to "St. Paul", so they get higher scores than the longer "St Paul" schools with extra suffixes (like "St Paul's Co-educational College").

Step 1: Standardize Text with Preprocessing

First, we'll clean up both the search query and school list to eliminate punctuation inconsistencies:

from fuzzywuzzy import process, fuzz

def preprocess_text(text):
    # Remove dots and trim any extra whitespace
    return text.replace(".", "").strip()

# Your original school list
schList = ["Diocesan Boy's School", "Diocesan Girl's School", 'Heep Yunn School', 'La Salle College', 'Maryknoll Convent School', 'Marymount Secondary School', 'Methodist College', 'Sacred Heart Canossian College', "St Clare's Girl's School", 'St Francis Canossian College', "St Joseph's College", "St Mark's School", "St Mary's Canossian College", "St Paul's Co-educational College", "St Paul's College", "St Paul's Convent School", "St Paul's Secondary School", "St Stephen's Girl's College", 'Wah Yan College, Hong Kong', 'Wah Yan College, Kowloon', 'Ying Wa College', "Ying Wa Girl's School"]

# Preprocess the school list and query
cleaned_schools = [preprocess_text(school) for school in schList]
cleaned_query = preprocess_text("St. Paul")

Step 2: Use a Scorer That Fits Your Needs

If you want to prioritize schools that start with (or contain) "St Paul", fuzz.partial_ratio is perfect—it focuses on partial matches instead of full-string similarity, so the suffixes on the "St Paul" schools won't drag down their scores.

Here's how to implement it:

# Get matching results with partial ratio scoring
matches = process.extractBests(cleaned_query, cleaned_schools, scorer=fuzz.partial_ratio, score_cutoff=85)

# Map back to the original school names (since we cleaned them)
final_results = [(schList[cleaned_schools.index(cleaned_school)], score) for cleaned_school, score in matches]

print(final_results)

This should return exactly the 4 schools you're expecting, with high scores (usually 90+).

If you prefer full-string matching, even just preprocessing will boost the scores of the "St Paul" schools enough to push them to the top:

matches = process.extractBests(cleaned_query, cleaned_schools, score_cutoff=80)
final_results = [(schList[cleaned_schools.index(cleaned_school)], score) for cleaned_school, score in matches]
print(final_results)

Bonus: Tweak Limits and Cutoffs

  • Use the limit parameter to control how many results you get (e.g., limit=4 to only return the top 4 matches).
  • Adjust score_cutoff to filter out low-quality matches entirely.

内容的提问来源于stack exchange,提问作者JOSEPH CHONG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 11:43:12