正则多字符串替换的词边界指定及性能优化问题
问题解决:城市名掩码的误匹配与性能优化
1. 误匹配问题修复
误匹配的核心原因是正则表达式未添加词边界,导致城市名子串被错误识别。解决方法:
- 对每个城市名添加词边界(\b),确保仅匹配完整的城市名短语
- 对城市名进行正则转义(re.escape),避免特殊字符干扰正则引擎
- 按城市名长度降序排序,优先匹配更长的城市名(避免短名被先匹配)
2. 性能优化
处理5000个城市时速度慢的原因:
- 每次调用函数都重复生成城市列表
- 未排序的正则备选项导致引擎回溯过多
优化措施:
- 预计算多词城市列表,避免重复处理
- 按城市名长度降序排序正则备选项,减少回溯
修正后的代码
import re import difflib # 预计算多词城市列表(仅执行一次) us_cities_all = ['Great Barrington', 'Round O', 'East Orange'] # 示例数据 multi_word_cities = list(set([ city for city in us_cities_all if len(city.split(' ')) > 1 and len(city) > 3 and "Mc" not in city and "State" not in city and city != 'Mary D' ])) def sub_mult_regex(text, keys, tag_type): ''' Replaces/masks multiple words at once Parameters: Text: TIU note Keys: a list of words to be replaced by the regex Tag_type: string you want the words to be replaced with Creates a replacement dictionary of keys and values (values are the length of the key, preserving formatting). Eg., {68 Oak St., PAddress PAddress PAddress.,} Returns text with relevant text masked ''' # 创建替换值列表:保留标点格式,仅替换单词部分为标签 add_vals = [] for val in keys: add_vals.append(re.sub(r'\w{1,100}', tag_type, val)) # 构建键值对字典 add_dict = dict(zip(keys, add_vals)) # 按长度降序排序城市名,优先匹配更长的短语 sorted_keys = sorted(add_dict.keys(), key=lambda x: -len(x)) # 编译正则:添加词边界和转义,避免误匹配 pattern = "|".join(f"(\\b{re.escape(key)}\\b)" for key in sorted_keys) add_subs = re.compile(pattern, re.IGNORECASE) # 构建索引化替换字典(与正则分组顺序一致) group_index = 1 indexed_subs = {} for key in sorted_keys: indexed_subs[group_index] = add_dict[key] group_index += re.compile(re.escape(key)).groups + 1 # 执行替换 if len(indexed_subs) > 0: text_sub = re.sub(add_subs, lambda match: indexed_subs[match.lastindex], text) else: text_sub = text # 生成差异列表 case_a = text case_b = text_sub diff_list = [li for li in difflib.ndiff(case_a.split(), case_b.split()) if li[0] != ' '] diff_list = [re.sub(r'[-,]', "", term.strip()) for term in diff_list if '-' in term] return text_sub, diff_list def mask_multiword_cities(text_string): return sub_mult_regex(text_string, multi_word_cities, "PAddress") # 测试 add_string = "The cities are Round O , NJ and around others" result = mask_multiword_cities(add_string) print(result) # 输出:('The cities are PAddress PAddress , NJ and around others', ['Round', 'O'])
关键说明
- 词边界与转义:
\b确保城市名作为完整单词匹配,re.escape处理城市名中可能含有的正则特殊字符(如点号、括号) - 降序排序:避免短城市名被优先匹配(例如先匹配"New York"而非"York")
- 预计算列表:减少重复处理
us_cities_all的开销,大幅提升多次调用的速度
内容的提问来源于stack exchange,提问作者skaleidoscope
相关产品推荐
相关产品推荐

