使用fuzzywuzzy的Process.extract获取多查询的所有最高相似度匹配项
Solution Using fuzzywuzzy's process.extract()
Got it, let's walk through exactly how to solve this problem. We'll use fuzzywuzzy's process.extract() method to fetch the highest similarity matches for each query, and make sure we keep all ties when multiple options have the same top score.
Step 1: Install fuzzywuzzy
First, make sure you have the library installed (adding python-Levenshtein is optional but speeds up the similarity calculations):
pip install fuzzywuzzy python-Levenshtein
Step 2: Code Implementation
Here's the full code to achieve your desired output. We'll loop through each query, grab all similarity matches, filter for the highest-scoring ones, and output the results:
from fuzzywuzzy import process # Define your input queries and choices queries = ('Hello', 'bye', 'yellow') choices = ('Hello', 'can', 'sweet', 'lye', 'fellow', 'Hello') # Process each query to get top matches for query in queries: # Get all matches with their similarity scores (out of 100) all_matches = process.extract(query, choices) # Find the highest score from the matches max_score = max(score for _, score in all_matches) # Filter to keep only matches that hit the maximum score top_matches = [match for match in all_matches if match[1] == max_score] # Print the result (you can also return this as a structured list if needed) print(f"Query: '{query}' → Top matches: {top_matches}")
Step 3: Expected Output
When you run the code, you'll get exactly the results you're looking for:
- Query: 'Hello' → Top matches: [('Hello', 100), ('Hello', 100)]
- Query: 'bye' → Top matches: [('lye', 80)]
- Query: 'yellow' → Top matches: [('fellow', 80)]
Quick Breakdown
process.extract()returns a list of tuples, each pairing an option fromchoiceswith its similarity score compared to the query.- We first find the highest score in that list, then filter out all matches that don't have that score—this ensures we keep every option that ties for first place (like the two 'Hello' entries here).
- The scoring uses the Levenshtein distance algorithm by default, which counts the number of character edits (additions, deletions, swaps) needed to turn one string into another. That's why 'bye' matches 'lye' (one character swap) and 'yellow' matches 'fellow' (shared character structure and partial overlap) both get an 80% score.
内容的提问来源于stack exchange,提问作者Lokesh Choraria
相关产品推荐
相关产品推荐

