正则表达式提取街道名及可选房屋字母时去除多余空格的问题
Great question! The issue with your current regex ([A-z ]+) is twofold: first, [A-z] actually includes some non-alphabet characters (like [, \, etc., since it’s based on ASCII order), and second, it greedily matches all letters and spaces—so it grabs those trailing/leading spaces you don’t want. Let’s fix this with targeted patterns tailored to your address structure.
Solution 1: Single regex to capture both targets in one go
Since your address follows the pattern [Street Name] [Number] [Letter] [Number], we can craft a regex that locks onto this structure, avoiding extra spaces entirely:
^([A-Za-z ]+?)\s*\d+\s*([A-Za-z])\s*\d+$
^([A-Za-z ]+?): The non-greedy+?matches the street name (letters and spaces) but stops as soon as it hits the next part—this avoids capturing trailing spaces before the number.\s*\d+: Matches optional spaces followed by the first number (103 in your case).\s*([A-Za-z])\s*: Captures the single letter (B) while ignoring any surrounding spaces.\d+$: Matches the final number (30) to anchor the end of the string.
In code (using Python as an example):
import re address = "Castle street 103 B 30" match = re.match(r'^([A-Za-z ]+?)\s*\d+\s*([A-Za-z])\s*\d+$', address) if match: street = match.group(1) # Output: "Castle street" unit = match.group(2) # Output: "B"
Solution 2: Separate regexes for each target
If you prefer to extract each part independently:
- For the street name:
^([A-Za-z ]+?)\s*\d+(captures everything up to the first number, no trailing space) - For the letter:
\d+\s*([A-Za-z])\s*\d+(captures the letter between two numbers, no surrounding spaces)
Alternative: Clean up captured text with string methods
If you want to stick closer to your original regex, you can use .strip() to remove leading/trailing spaces from the captured groups:
import re address = "Castle street 103 B 30" # Capture street name (with possible trailing space) street_match = re.search(r'([A-Za-z ]+)', address) street = street_match.group(1).strip() # "Castle street" # Capture the letter (with possible surrounding spaces) unit_match = re.search(r'([A-Za-z])', address[address.index('103'):]) unit = unit_match.group(1).strip() # "B"
This is less precise than the targeted regex, but works if your address format varies slightly.
内容的提问来源于stack exchange,提问作者melvinb

