求助:基于近似匹配对两组元组列表进行排序
Hey, I get exactly what you're dealing with here—strict exact matches just aren't cutting it for those edge cases like trailing spaces, abbreviations (looking at you, Nth vs North Queensland Cowboys), right? Let's fix this with a combination of string preprocessing and fuzzy similarity matching—it's way more flexible and fits your 85% similarity threshold need perfectly.
First: Clean up the team names (preprocessing)
Before we even start matching, let's standardize the team names to eliminate trivial differences:
- Trim leading/trailing spaces (like that extra space in
Canterbury Bulldogs) - Convert everything to lowercase to avoid case sensitivity issues
- Map common abbreviations to their full forms (like
Nth→North,St.→St)
Second: Use fuzzy matching for similarity checks
Python has a super handy library called fuzzywuzzy that calculates string similarity using Levenshtein distance (how many edits it takes to turn one string into another). It returns a score from 0-100, which makes it easy to set your 85% threshold.
Install the dependencies
First, grab the library (the second package speeds things up, but it's optional):
pip install fuzzywuzzy python-Levenshtein
Third: Put it all together with code
Here's a complete replacement for your original matching logic, with preprocessing and fuzzy checks built in:
from fuzzywuzzy import fuzz # Preprocess team names to standardize them def preprocess_name(name): # Trim spaces, lowercase, fix common abbreviations cleaned = name.strip().lower() cleaned = cleaned.replace("nth", "north") cleaned = cleaned.replace("st.", "st") # Add other abbreviation fixes here if you run into more cases return cleaned # Check if two names are a match based on similarity threshold def is_match(name1, name2, threshold=85): return fuzz.ratio(preprocess_name(name1), preprocess_name(name2)) >= threshold # Your original data (unchanged) list_finale = [[('Canterbury Bulldogs ', '3.25'), ('South Sydney Rabbitohs', '1.34')], [('Parramatta Eels ', '1.79'), ('Wests Tigers', '2.02')], [('Melbourne Storm ', '1.90'), ('Sydney Roosters', '1.90')], [('Gold Coast Titans ', '1.86'), ('Newcastle Knights', '1.94')], [('New Zealand Warriors ', '1.39'), ('North Queensland Cowboys', '2.95')], [('Cronulla Sharks ', '1.68'), ('Penrith Panthers', '2.18')], [('St. George Illawarra Dragons ', '1.45'), ('Manly Sea Eagles', '2.74')], [('Canberra Raiders ', '1.63'), ('Brisbane Broncos', '2.26')]] sportsbet_list = [[('Cronulla Sharks', '1.64'), ('Penrith Panthers', '2.27')], [('Canterbury Bulldogs', '3.30'), ('South Sydney Rabbitohs', '1.33')], [('Melbourne Storm', '1.90'), ('Sydney Roosters', '1.90')], [('New Zealand Warriors', '1.40'), ('Nth Queensland Cowboys', '2.90')], [('St George Illawarra Dragons', '1.45'), ('Manly Sea Eagles', '2.75')], [('Gold Coast Titans', '1.85'), ('Newcastle Knights', '1.95')], [('Canberra Raiders', '1.60'), ('Brisbane Broncos', '2.30')], [('Parramatta Eels', '1.90'), ('Wests Tigers', '1.90')], [('Sydney Roosters', '1.35'), ('St George Illawarra Dragons', '3.20')], [('Melbourne Storm', '1.25'), ('New Zealand Warriors', '4.00')], [('Canterbury Bulldogs', '1.56'), ('Nth Queensland Cowboys', '2.40')], [('Penrith Panthers', '2.20'), ('South Sydney Rabbitohs', '1.67')], [('Wests Tigers', '1.67'), ('Gold Coast Titans', '2.20')], [('Brisbane Broncos', '1.70'), ('Cronulla Sharks', '2.15')], [('Manly Sea Eagles', '1.85'), ('Canberra Raiders', '1.95')], [('Newcastle Knights', '1.80'), ('Parramatta Eels', '2.00')]] list_n = [] list_n1 = [] # Loop through both lists to find fuzzy matches for sportsbet_match in sportsbet_list: for finale_match in list_finale: # Check both possible team orderings (A vs B or B vs A) order_match = is_match(sportsbet_match[0][0], finale_match[0][0]) and is_match(sportsbet_match[1][0], finale_match[1][0]) reverse_order_match = is_match(sportsbet_match[0][0], finale_match[1][0]) and is_match(sportsbet_match[1][0], finale_match[0][0]) if order_match or reverse_order_match: list_n.append(sportsbet_match) list_n1.append(finale_match) break # Stop checking other finale matches once we find a hit # Print results to verify for sb, fn in zip(list_n, list_n1): print(f"Sportsbet: {sb}") print(f"Matched Finale: {fn}") print("---")
Quick notes to tweak for your use case
- Adjust the threshold: If you're getting too many false matches, bump the threshold up (like 90). If you're missing valid matches, lower it a bit (like 80).
- Add more preprocessing rules: If you run into other odd formatting (like hyphens vs spaces, or nicknames), add them to the
preprocess_namefunction. - Skip reverse order check: If you know the team order in both lists is always the same, you can remove the
reverse_order_matchcheck to speed things up. - No third-party library?: If you can't install
fuzzywuzzy, you can implement your own Levenshtein distance function, but the library is way easier for someone new to coding.
内容的提问来源于stack exchange,提问作者Jordan

