循环中将变量整合进复杂正则失效,求正则优化方案
问题与解决方案
问题描述
我想编写代码检测字符串中城市关键词后续的若干单词,判断其中是否包含州名。单独用硬编码的城市名(比如"Orange")测试正则能正常匹配,但将城市名放到循环里用变量传入时,正则完全匹配不到结果,试过re.escape()处理城市变量也无效。此外,希望优化正则,使其能匹配城市后的前三个单词,无论城市与这些单词之间存在标点、换行还是空白字符——例如在字符串"123 Cherry Lane , Orange , \n \n \n 13238727 New Jersey is"中,能提取出"New Jersey is"。
失效的循环代码
import re cities = ["Orange", "Denver"] test_str = "The city is Orange, NJ, 07052" for city in cities: regex = r"(?<=\b" + city + r")\s?(,|\.|\s){1,100}(\w{1,100})?\s?[,.\-]?\s{1,100}?(\w{1,100})?\s{1,100}?(\w{1,100})?" matches = re.finditer(regex, test_str, re.MULTILINE) for matchNum, match in enumerate(matches, start=1): print("Match {matchNum} for {city} was found at {start}-{end}: {match}".format(matchNum=matchNum, city=city, start=match.start(), end=match.end(), match=match.group())) for groupNum in range(0, len(match.groups())): groupNum = groupNum + 1 print("Group {groupNum} found at {start}-{end}: {group}".format(groupNum=groupNum, start=match.start(groupNum), end=match.end(groupNum), group=match.group(groupNum)))
硬编码生效的代码
regex = r"(?<=\bOrange)\s?(,|\.|\s){1,100}(\w{1,100})?\s?[,.\-]?\s{1,100}?(\w{1,100})?\s{1,100}?(\w{1,100})?" test_str = "The city is Orange, New Jersey" matches = re.finditer(regex, test_str, re.MULTILINE) for matchNum, match in enumerate(matches, start=1): print ("Match {matchNum} was found at {start}-{end}: {match}".format(matchNum = matchNum, start = match.start(), end = match.end(), match = match.group())) for groupNum in range(0, len(match.groups())): groupNum = groupNum + 1 print ("Group {groupNum} found at {start}-{end}: {group}".format(groupNum = groupNum, start = match.start(groupNum), end = match.end(groupNum), group = match.group(groupNum)))
问题原因
- 正则规则冗余复杂:原正则使用过多可选分组和重复的标点/空白匹配规则,导致贪婪匹配逻辑混乱,容易出现匹配不到有效内容的情况。
- 匹配逻辑冲突:原正则的后行断言本身没问题,但后续的可选分组可能导致正则匹配到空内容,而
re.finditer不会输出空匹配结果。
解决方案
优化后的正则与代码
import re cities = ["Orange", "Denver"] test_str = "The city is Orange, NJ, 07052\n123 Cherry Lane , Orange , \n \n \n 13238727 New Jersey is\nDenver, Colorado" for city in cities: # 编译正则:匹配完整城市单词,后续任意非单词字符,再提取最多3个词 regex = re.compile(rf"\b{re.escape(city)}\b\W+(?:\W*\w+){1,3}", re.MULTILINE | re.DOTALL) matches = regex.finditer(test_str) for match_num, match in enumerate(matches, start=1): print(f"Match {match_num} for {city} was found at {match.start()}-{match.end()}: {match.group()}") # 提取城市后的单词部分 post_city_content = match.group().replace(city, '', 1).strip() # 分割为单词列表(过滤空值) extracted_words = [word for word in re.split(r'\W+', post_city_content) if word] print(f"提取到的城市后单词(最多3个):{extracted_words[:3]}\n")
代码说明
re.escape(city):自动转义城市名中的特殊字符(如.、+等),避免被当作正则语法解析。\b{re.escape(city)}\b:确保匹配完整的城市单词,避免部分匹配(如"Orange"不会匹配"Orangeville")。\W+:匹配城市后的任意数量非单词字符(包括标点、空格、换行符等)。(?:\W*\w+){1,3}:匹配1-3个单词,每个单词前允许有任意非单词字符,非捕获组(?:...)避免生成多余分组。re.DOTALL:让正则能匹配换行符,确保跨换行的内容也能被识别。
内容的提问来源于stack exchange,提问作者skaleidoscope
相关产品推荐
相关产品推荐

