如何从完整地址数组提取城市并实现职位按城市筛选?
Hey there, let's tackle this messy address filtering problem you're facing—those unstructured scraped addresses and tricky city aliases (like Lefkosia/Nicosia) can be a pain, but we can break this down into manageable steps.
Step 1: Standardize Address Data First
Your raw addresses are inconsistent, so the first move is to extract clean, consistent city names from each entry. The key here is to leverage patterns in the data (like country codes, postal codes, and capitalization) with regex, tailored to common European address formats.
Example Regex Rules by Country
Looking at your sample data, we can create country-specific regex patterns to pull out cities:
- Cyprus (cy): Addresses start with
[CITY] [POSTAL CODE] cyor end with[POSTAL CODE] [CITY](all caps). Use regex like^([A-Z\s]+)\s+\d{4}\s+cyto grab the starting city, or\d{4}\s+([A-Z\s]+)$for the trailing city. - Belgium (be): Addresses start with
[CITY] [POSTAL CODE] be—regex^([A-Z\s]+)\s+\d{4}\s+beworks here. - Other European Countries: Extend this logic:
- Germany (de): Addresses end with
[POSTAL CODE] [CITY]→ use\d{5}\s+([A-Za-z\s]+)$. - France (fr): Addresses include
[POSTAL CODE] [CITY]→ use\d{5}\s+([A-Za-z\s]+).
- Germany (de): Addresses end with
Step 2: Handle City Aliases with a Mapping Table
To fix the Lefkosia/Nicosia problem (and other European city aliases), create a centralized alias map. This lets you group equivalent cities together so selecting either name pulls up the same results.
Example Alias Map
city_alias_map = { "LEFKOSIA": {"NICOSIA"}, "NICOSIA": {"LEFKOSIA"}, # Add more as needed for other countries: # "DEN HAAG": {"THE HAGUE"}, # "THE HAGUE": {"DEN HAAG"}, # "MILANO": {"MILAN"}, # "MILAN": {"MILANO"} }
When a user selects a city, expand the search to include all its aliases (normalized to uppercase/lowercase to avoid case sensitivity issues).
Step 3: Build the Filtering Logic
Combine the address extraction and alias mapping to create your filtering workflow:
- Normalize the user's selected city (e.g., convert to uppercase).
- Fetch all aliases for that city from your map.
- Extract the city from each job's address using the country-specific regex.
- Match the extracted city against the selected city + aliases.
Full Example Code (Python)
import re # Centralized alias map (extend this as you add countries) city_alias_map = { "LEFKOSIA": {"NICOSIA"}, "NICOSIA": {"LEFKOSIA"} } def extract_city_from_address(address_str): """Extract clean city name from unstructured address string""" # First, get the 2-letter country code (e.g., cy, be) country_match = re.search(r'\b([a-z]{2})\b', address_str.lower()) if not country_match: return None country_code = country_match.group(1) # Cyprus-specific extraction if country_code == "cy": # Try extracting starting city first start_city_match = re.search(r'^([A-Z\s]+)\s+\d{4}\s+cy', address_str) if start_city_match: return start_city_match.group(1).strip() # Fallback to trailing city end_city_match = re.search(r'\d{4}\s+([A-Z\s]+)$', address_str) if end_city_match: return end_city_match.group(1).strip() # Belgium-specific extraction elif country_code == "be": city_match = re.search(r'^([A-Z\s]+)\s+\d{4}\s+be', address_str) if city_match: return city_match.group(1).strip() # Add more country rules here (Germany, France, etc.) return None def get_matching_city_set(selected_city): """Get the set of cities including the selected one and its aliases""" selected_upper = selected_city.strip().upper() matching_cities = {selected_upper} # Add aliases if they exist if selected_upper in city_alias_map: matching_cities.update(city_alias_map[selected_upper]) return matching_cities # Test with your sample addresses sample_addresses = [ "LEMESOS 3042 cy ΡΙΧΑΡΔΟΥ & ΒΕΡΕΓΓΑΡΙΑΣ 12 ARAOUZOS CASTLE COURT, 3ΟΣ ΟΡΟΦΟΣ 3042 LEMESOS", "LARNAKA - SOTIR 6057 cy ΣΠΥΡΟΥ ΚΥΠΡΙΑΝΟΥ 50, ΙΡΙΔΑ 3 10-ΟΣ ΟΡΟΦΟΣ 6057 LARNAKA", "EDEGEM 2650 be Acht Eeuwenlaan", "STROVOLOS 2064 cy ΒΥΖΑΝΤΙΟΥ 30 FLAT/ OFFICE 22 2064 LEFKOSIA", "LEMESOS 3042 cy ΡΙΧΑΡΔΟΥ & ΒΕΡΕΓΓΑΡΙΑΣ 12 ARAOUZOS CASTLE COURT, 3ΟΣ ΟΡΟΦΟΣ 3042 LEMESOS", "LEMESOS 3042 cy ΡΙΧΑΡΔΟΥ & ΒΕΡΕΓΓΑΡΙΑΣ 12 ARAOUZOS CASTLE COURT, 3ΟΣ ΟΡΟΦΟΣ 3042 LEMESOS", "EGKOMI 2404 cy Διογένους 1, Κόμβος A, 5ος όροφος 2404 LEFKOSIA", "EGKOMI 2404 cy Διογένους 1, Κόμβος A, 5ος όροφος 2404 LEFKOSIA", "LAKATAMEIA 2322 cy Arch. Makariou III and Mesaorias 1 2322 LEFKOSIA" ] # Simulate user selecting "Nicosia" user_selected_city = "Nicosia" matching_cities = get_matching_city_set(user_selected_city) # Filter addresses matching_jobs = [] for addr in sample_addresses: extracted_city = extract_city_from_address(addr) if extracted_city and extracted_city in matching_cities: matching_jobs.append(addr) print(f"Found {len(matching_jobs)} jobs for '{user_selected_city}':") for job in matching_jobs: print(job)
Running this will return the 4 Lefkosia/Nicosia entries you mentioned—perfect!
Step 4: Optimize for Scalability & Accuracy
- Use Address Standardization Tools: If you have the budget, tools like Google Maps Geocoding or OpenStreetMap's Nominatim can automatically parse messy addresses into structured data (city, postal code, country) with high accuracy, saving you from maintaining regex rules for every country.
- Maintain a Dynamic Alias Database: Store aliases in a database or config file instead of hardcoding—this makes it easy to add new aliases as you support more countries.
- Add User Input Tolerance: Normalize user input (lowercase/uppercase stripping, typo correction with fuzzy matching) to handle cases where users enter "nicosia" instead of "Nicosia" or minor typos.
- Test with Diverse Samples: Collect address samples from all the European countries you want to support and validate your extraction logic against them to catch edge cases.
内容的提问来源于stack exchange,提问作者user14054150

