You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则多字符串替换的词边界指定及性能优化问题

问题解决:城市名掩码的误匹配与性能优化

1. 误匹配问题修复

误匹配的核心原因是正则表达式未添加词边界,导致城市名子串被错误识别。解决方法:

  • 对每个城市名添加词边界(\b),确保仅匹配完整的城市名短语
  • 对城市名进行正则转义(re.escape),避免特殊字符干扰正则引擎
  • 按城市名长度降序排序,优先匹配更长的城市名(避免短名被先匹配)

2. 性能优化

处理5000个城市时速度慢的原因:

  • 每次调用函数都重复生成城市列表
  • 未排序的正则备选项导致引擎回溯过多

优化措施:

  • 预计算多词城市列表,避免重复处理
  • 按城市名长度降序排序正则备选项,减少回溯

修正后的代码

import re
import difflib

# 预计算多词城市列表(仅执行一次)
us_cities_all = ['Great Barrington', 'Round O', 'East Orange']  # 示例数据
multi_word_cities = list(set([
    city for city in us_cities_all 
    if len(city.split(' ')) > 1 and len(city) > 3 
    and "Mc" not in city and "State" not in city and city != 'Mary D'
]))

def sub_mult_regex(text, keys, tag_type):
    '''
    Replaces/masks multiple words at once
    Parameters:
        Text: TIU note
        Keys: a list of words to be replaced by the regex
        Tag_type: string you want the words to be replaced with
    Creates a replacement dictionary of keys and values 
    (values are the length of the key, preserving formatting).
    Eg., {68 Oak St., PAddress PAddress PAddress.,}
    Returns text with relevant text masked
    '''
    # 创建替换值列表:保留标点格式,仅替换单词部分为标签
    add_vals = []
    for val in keys:
        add_vals.append(re.sub(r'\w{1,100}', tag_type, val))
    
    # 构建键值对字典
    add_dict = dict(zip(keys, add_vals))
    
    # 按长度降序排序城市名,优先匹配更长的短语
    sorted_keys = sorted(add_dict.keys(), key=lambda x: -len(x))
    
    # 编译正则:添加词边界和转义,避免误匹配
    pattern = "|".join(f"(\\b{re.escape(key)}\\b)" for key in sorted_keys)
    add_subs = re.compile(pattern, re.IGNORECASE)
    
    # 构建索引化替换字典(与正则分组顺序一致)
    group_index = 1
    indexed_subs = {}
    for key in sorted_keys:
        indexed_subs[group_index] = add_dict[key]
        group_index += re.compile(re.escape(key)).groups + 1
    
    # 执行替换
    if len(indexed_subs) > 0:
        text_sub = re.sub(add_subs, lambda match: indexed_subs[match.lastindex], text)
    else:
        text_sub = text
    
    # 生成差异列表
    case_a = text
    case_b = text_sub
    diff_list = [li for li in difflib.ndiff(case_a.split(), case_b.split()) if li[0] != ' ']
    diff_list = [re.sub(r'[-,]', "", term.strip()) for term in diff_list if '-' in term]
    
    return text_sub, diff_list 

def mask_multiword_cities(text_string):
    return sub_mult_regex(text_string, multi_word_cities, "PAddress")

# 测试
add_string = "The cities are Round O , NJ and around others"
result = mask_multiword_cities(add_string)
print(result)
# 输出:('The cities are PAddress PAddress , NJ and around others', ['Round', 'O'])

关键说明

  • 词边界与转义:\b确保城市名作为完整单词匹配,re.escape处理城市名中可能含有的正则特殊字符(如点号、括号)
  • 降序排序:避免短城市名被优先匹配(例如先匹配"New York"而非"York")
  • 预计算列表:减少重复处理us_cities_all的开销,大幅提升多次调用的速度

内容的提问来源于stack exchange,提问作者skaleidoscope

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 05:20:08