从英国格式地址字符串中提取门牌号的技术实现问询
Hey there! Let's dive into your UK address house number extraction code, break down how it works, highlight its strengths, and go over some potential improvements and edge cases you might want to handle.
How Your Current Code Works
First, let's recap the core logic:
- It uses
re.search('\d+', address)to locate the starting position of the first sequence of digits in the address string. - If those digits sit at the very start of the string (
position == 0), it loops through characters until hitting a space, capturing everything before that space as the house number (perfect for cases like'6A McCarthy Way'). - If the digits aren't at the start, it assumes the house number is everything from that digit position to the end of the string (handles reverse cases like
'McCarthy Way 6A').
Strengths of Your Implementation
- It directly targets the two most common UK address layouts: house number leading, or house number trailing the street name.
- The core logic is straightforward and easy to follow at a glance.
Potential Edge Cases & Limitations
There are a few scenarios where your code might run into issues:
- Addresses with extra numbers (e.g.,
'Flat 2, 6A McCarthy Way'): Your code will pick up the2first, and since it's not at position 0, it'll return'2, 6A McCarthy Way'—which is incorrect, as the actual house number is6A. - Addresses without digits (e.g.,
'The Cottage, Main Road'):re.search('\d+', address)will returnNone, triggering anAttributeErrorwhen you try to access.start(). - Redundant loop: The
for i in range(0,100000)loop is overkill—you don't need to check 100k characters when you can just find the first space directly. - Special characters in house numbers: While your code handles letters after digits (like
6A), it doesn't explicitly account for common UK house number symbols like hyphens (12-A) or slashes (10/12)—though it does work for these cases since it stops only at spaces.
Optimized Implementation
We can refine this with more targeted regex patterns that better align with UK house number conventions, while fixing the above limitations:
import re def extract_house_number(address): # Match house number at the start (followed by a space) start_match = re.match(r'^([\d\w/-]+)\s', address) if start_match: return start_match.group(1) # Match house number at the end (preceded by a space) end_match = re.search(r'\s([\d\w/-]+)$', address) if end_match: return end_match.group(1) # Fallback: Find the first sequence containing digits (handles odd placements) fallback_match = re.search(r'([\d]+[\w/-]*|[\w/-]*[\d]+)', address) if fallback_match: return fallback_match.group(1) # No house number found return None
Key Improvements Here:
- Precise regex patterns: The patterns explicitly match common UK house number characters (digits, letters, hyphens, slashes) in the typical start/end positions.
- Error prevention: It avoids crashes when no digits are present by checking for match existence before accessing groups.
- Efficient logic: No unnecessary loops—regex handles the pattern matching in one pass.
- Fallback handling: Catches edge cases where the house number might be in an unusual position (though this is rare for UK addresses).
Quick Fix for Your Original Code
If you want to keep your core logic but fix the redundant loop, replace the loop section with this:
position_match = re.search('\d+', address) if position_match: position = position_match.start() if position == 0: space_pos = address.find(' ') house_number = address[:space_pos] if space_pos != -1 else address else: house_number = address[position:] else: # Handle addresses with no house number house_number = None
This replaces the 100k iteration loop with a single find() call, which is far more efficient.
内容的提问来源于stack exchange,提问作者LucSpan
相关产品推荐
相关产品推荐

