You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

修改Python2.7正则识别短/数字街道名 避免AttributeError报错

Python 2.7 街名正则匹配脚本优化

问题说明

原有Python 2.7环境下的正则匹配脚本,用于识别字符串中的street(街道)、avenue(大道)名称,存在两类问题:

  • 匹配覆盖范围不足:仅能匹配Stackoverflow Street/Stackoverflow Avenue这类常规名称,无法识别:
    • 仅含2个字母的极短街名,例如SO Street、SO Avenue
    • 带数字序号的街名,例如29th Street、1st Avenue
  • 逻辑缺陷:正则未匹配到结果时,直接对None返回值调用group()方法,会触发AttributeError: 'NoneType' object has no attribute 'group'致命错误,导致程序终止。

原有可运行代码

import re

def street(search):
    if bool(re.search('(?i)street', search)):
        found = re.search('([A-Z]\S[a-z]+\s(?i)street)', search)
        found = found.group()
        return found.title()
    if bool(re.search('(?i)avenue', search)):
        found = re.search('([A-Z]\S[a-z]+\s(?i)avenue)', search)
        found = found.group()
        return found.title()
    else:
        found = "na"
        return found

userlocation = street("I live on Stackoverflow Street")

print userlocation

失败测试用例

以下测试用例运行时无法得到预期结果:

userlocation = street("I live on SO Street")
userlocation = street("I live on 29th Street")
userlocation = street("I live on 1st Avenue")
userlocation = street("I live on SO Avenue")

运行报错信息

me@me:~/Documents/test$ python2.7 test_street.py
Traceback (most recent call last):
  File "test_street.py", line 12, in <module>
    userlocation = street("I live on 29th Street")
  File "test_street.py", line 6, in street
    found = found.group()
AttributeError: 'NoneType' object has no attribute 'group'

优化实现

优化点说明

  • 正则规则调整:重构匹配逻辑,同时支持三类街名前缀:常规首字母大写的单词街名、2位及以上大写字母的极短街名、数字+st/nd/rd/th格式的数字序号街名;合并street和avenue的匹配逻辑,减少重复正则判断。
  • 容错逻辑优化:正则匹配后先判断返回结果是否为空,确认匹配成功后再调用group()方法,避免空值触发属性错误。

优化后完整代码

import re

def street(search):
    # 正则规则说明:
    # 前缀支持三类:首字母大写的单词、2位以上全大写短名、数字+序号后缀(st/nd/rd/th)
    # 后缀匹配street/avenue,开启不区分大小写flag
    match_result = re.search(
        r'((?:[A-Z][a-zA-Z]+|[A-Z]{2,}|\d+(?:st|nd|rd|th))\s(?:street|avenue))',
        search,
        flags=re.IGNORECASE
    )
    if match_result:
        return match_result.group().title()
    return "na"

if __name__ == "__main__":
    # 全量测试用例
    test_cases = [
        "I live on Stackoverflow Street",
        "I live on SO Street",
        "I live on 29th Street",
        "I live on 1st Avenue",
        "I live on SO Avenue",
        "Random text with no address"
    ]
    for case in test_cases:
        print street(case)

运行结果

Stackoverflow Street
So Street
29Th Street
1St Avenue
So Avenue
na

注:因原代码使用title()方法格式化返回结果,序号后缀的st/nd/rd/th会被转为首字母大写格式,若需保留全小写的序号格式,可单独对序号部分做字符串处理,不影响核心匹配逻辑。

内容的提问来源于stack exchange,提问作者ratuk_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 04:06:06