You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式提取街道名及可选房屋字母时去除多余空格的问题

Great question! The issue with your current regex ([A-z ]+) is twofold: first, [A-z] actually includes some non-alphabet characters (like [, \, etc., since it’s based on ASCII order), and second, it greedily matches all letters and spaces—so it grabs those trailing/leading spaces you don’t want. Let’s fix this with targeted patterns tailored to your address structure.

Solution 1: Single regex to capture both targets in one go

Since your address follows the pattern [Street Name] [Number] [Letter] [Number], we can craft a regex that locks onto this structure, avoiding extra spaces entirely:

^([A-Za-z ]+?)\s*\d+\s*([A-Za-z])\s*\d+$
  • ^([A-Za-z ]+?): The non-greedy +? matches the street name (letters and spaces) but stops as soon as it hits the next part—this avoids capturing trailing spaces before the number.
  • \s*\d+: Matches optional spaces followed by the first number (103 in your case).
  • \s*([A-Za-z])\s*: Captures the single letter (B) while ignoring any surrounding spaces.
  • \d+$: Matches the final number (30) to anchor the end of the string.

In code (using Python as an example):

import re
address = "Castle street 103 B 30"
match = re.match(r'^([A-Za-z ]+?)\s*\d+\s*([A-Za-z])\s*\d+$', address)
if match:
    street = match.group(1)  # Output: "Castle street"
    unit = match.group(2)    # Output: "B"

Solution 2: Separate regexes for each target

If you prefer to extract each part independently:

  • For the street name: ^([A-Za-z ]+?)\s*\d+ (captures everything up to the first number, no trailing space)
  • For the letter: \d+\s*([A-Za-z])\s*\d+ (captures the letter between two numbers, no surrounding spaces)

Alternative: Clean up captured text with string methods

If you want to stick closer to your original regex, you can use .strip() to remove leading/trailing spaces from the captured groups:

import re
address = "Castle street 103 B 30"
# Capture street name (with possible trailing space)
street_match = re.search(r'([A-Za-z ]+)', address)
street = street_match.group(1).strip()  # "Castle street"
# Capture the letter (with possible surrounding spaces)
unit_match = re.search(r'([A-Za-z])', address[address.index('103'):])
unit = unit_match.group(1).strip()  # "B"

This is less precise than the targeted regex, but works if your address format varies slightly.

内容的提问来源于stack exchange,提问作者melvinb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:14:22