Python如何使用字典替换字符串中的复合词
复合词批量替换函数实现问题
问题背景
现有一个字典,键为文本中待替换的复合词,值为替换后的目标表达式,示例如下:
terms_dict = {'digi conso': 'digi conso', 'digi': 'digi conso', 'digiconso': 'digi conso', '3xcb': '3xcb', '3x cb': '3xcb', 'legal entity identifier': 'legal entity identifier'}
需要实现replace_terms(text, terms_dict)函数,接收文本和上述格式字典为入参,返回替换完成的文本。
测试用例:
test_text = "i want a digi conso loan for digiconso" print(replace_terms(test_text, terms_dict))
预期输出:
i want a digi conso loan for digi conso
最初尝试用.replace()方法实现,因待替换术语包含多词复合词无法正常生效。后续尝试基于正则多轮替换实现,代码如下:
def replace_terms(text, terms_dict): if len(terms_dict) > 0: words_in = [k for k in terms_dict.keys() if k in text] # ex: words_in = [digi conso, digi, digiconso] if len(words_in) > 0: for w in words_in: pattern = r"\b" + w + r"\b" text = re.sub(pattern, terms_dict[w], text) return text
该实现运行测试用例时返回*"i want a digi conso conso loan for digi conso"*,出现conso重复的问题。已知问题成因为:提前遍历字典键生成待匹配列表,替换过程中不会同步更新匹配状态,且短键会匹配长键替换后内容中的片段,触发重复替换。
实现方案
核心优化点有两个:
- 按待替换术语的长度倒序排序,优先匹配更长的复合词,避免短词拆分命中长词片段
- 构建统一正则模式做单轮扫描替换,避免多轮替换带来的连锁匹配问题
完整实现代码:
import re def replace_terms(text, terms_dict): if not terms_dict: return text # 待替换术语按包含单词数从多到少排序,长复合词优先匹配 sorted_terms = sorted(terms_dict.keys(), key=lambda x: len(x.split()), reverse=True) # 构建正则匹配模式,自动转义术语中的正则特殊字符 match_pattern = re.compile( r"\b(" + "|".join(re.escape(term) for term in sorted_terms) + r")\b" ) # 单轮扫描完成所有替换,匹配到的内容直接映射字典取值 return match_pattern.sub(lambda match_res: terms_dict[match_res.group(0)], text)
方案说明
运行测试用例可直接得到预期输出,该实现同时兼容术语包含特殊字符、多词嵌套等场景,单轮正则扫描的执行效率远高于多轮循环替换,适合大文本批量替换场景。
内容的提问来源于stack exchange,提问作者Perrupi
相关产品推荐
相关产品推荐

