Python中地址含国家/城市信息的高效搜索优化方案咨询
First, let’s fix a couple of critical issues in your current code that are breaking functionality before we even get to performance:
- The
combinationsvariable is an iterator, so it’ll be exhausted after the first address check—subsequent addresses won’t get any matches. - Your
returnstatements are inside the inner loop, meaning you only check the first combination for each address and immediately return, which is almost certainly not what you want.
Now, onto the performance fixes. Iterating through thousands of combinations per address is O(n*m) time complexity, which gets brutal fast as your datasets grow. Here are three far better approaches:
1. Precompiled Regular Expressions (Simplest & Effective for Most Cases)
Instead of generating every possible punctuation combination, build a single regex pattern that matches any target term (country code/name, city) surrounded by common separators or word boundaries. This lets you check an address in a single regex search.
import re # Example targets—replace with your full dataset from ISO 3166-1 and city lists target_terms = [ "GB", "GBR", "United Kingdom", "Cardiff", # Add all other country codes, names, and cities here ] # Escape special regex characters in each term (e.g., "Côte d'Ivoire" becomes safe) escaped_terms = [re.escape(term) for term in target_terms] # Build a pattern that matches the term surrounded by allowed punctuation or start/end of string # Accounts for spaces, commas, slashes, backslashes, semicolons, or edge of the address pattern = re.compile( r'(?:^|[\s,/\\;])(' + '|'.join(escaped_terms) + r')(?:$|[\s,/\\;])', re.IGNORECASE # Optional: make matching case-insensitive ) # Check addresses in a single pass per address def has_location_info(address): return pattern.search(address) is not None # Test with your example address test_address = "Steve; Studio 103, The Business Centre; 61 Wellfield Road; Cardiff; GB" print(has_location_info(test_address)) # Returns True
This approach reduces your per-address check to O(k) where k is the length of the address, instead of O(m) where m is the number of combinations.
2. Set-Based Token Matching (Fast for Exact Matches)
Split the address into "tokens" (split by punctuation) and check if any token exists in a pre-built set of target terms. Set lookups are O(1), making this extremely fast for most use cases.
import string # Preprocess target terms into a lowercase set for case-insensitive matching target_set = {term.lower() for term in target_terms} def has_location_info(address): # Replace all punctuation with spaces, then split into clean tokens translator = str.maketrans('', '', string.punctuation) cleaned_address = address.translate(translator).lower() tokens = cleaned_address.split() # Check for overlap between address tokens and target terms return bool(target_set.intersection(tokens)) # Test the example print(has_location_info(test_address)) # Returns True
This works best for single-token targets (like ISO codes or city names). For multi-word country names (e.g., "United States"), you can extend it to check for n-grams, but for most scenarios, this is blazingly fast.
3. Aho-Corasick Algorithm (Best for Large Pattern Sets)
If you have tens of thousands of target terms, the Aho-Corasick algorithm is designed for fast multi-pattern matching. It builds a trie of patterns and scans the address once to find all matches.
You’ll need the pyahocorasick library (install with pip install pyahocorasick):
import ahocorasick # Build the automaton A = ahocorasick.Automaton() # Add all target terms to the automaton (lowercase for case-insensitive matching) for idx, term in enumerate(target_terms): A.add_word(term.lower(), (idx, term)) A.make_automaton() def has_location_info(address): lowercase_addr = address.lower() # Scan the address for all matching terms for _, (_, term) in A.iter(lowercase_addr): # Verify the term is surrounded by valid separators to avoid partial matches match_pos = lowercase_addr.find(term) start_ok = match_pos == 0 or address[match_pos-1] in ' ,/\\;' end_ok = match_pos + len(term) == len(address) or address[match_pos+len(term)] in ' ,/\\;' if start_ok and end_ok: return True return False # Test the example print(has_location_info(test_address)) # Returns True
This is the most efficient option for very large datasets, running in O(k + z) time where k is the address length and z is the number of matches found.
内容的提问来源于stack exchange,提问作者M. Morgan

