如何在Python中使用正则替换句子列表中的多个子字符串
问题描述
我有如下句子列表:
sentences = ["I am learning to code", "coding seems to be intresting in python", "how to code in python", "practicing how to code is the key"]
现在我想通过存储了待替换字符串和对应替换值的字典,替换该句子列表中的若干子字符串:
word_list = {'intresting': 'interesting', 'how to code': 'learning how to code', 'am learning':'love learning', 'in python': 'using python'}
我尝试了如下代码:
replaced_sentences = [' '.join([word_list.get(w, w) for w in sentence.split()]) for sentence in sentences]
但运行后只有单个单词的字符串被替换,多单词的键没有生效。原因是我使用了sentence.split()按单词拆分句子,导致长度超过1个单词的子字符串没有被匹配到。
请问我要如何通过正则或其他方案实现子字符串的精确匹配替换?
预期输出
sentences = ["I love learning to code", "coding seems to be interesting using python", "learning how to code using python", "practicing learning how to code is the key"]
解决方案
你可以用正则匹配实现多词子串的替换,核心要注意优先匹配更长的键,避免短键先匹配破坏长键的完整结构,具体实现如下:
import re # 原始数据 sentences = ["I am learning to code", "coding seems to be intresting in python", "how to code in python", "practicing how to code is the key"] word_list = {'intresting': 'interesting', 'how to code': 'learning how to code', 'am learning':'love learning', 'in python': 'using python'} # 1. 按键的长度降序排序,保证长键优先匹配 sorted_keys = sorted(word_list.keys(), key=lambda x: len(x), reverse=True) # 2. 生成正则匹配模式,对每个键做转义处理避免特殊字符干扰 pattern = re.compile('|'.join(re.escape(key) for key in sorted_keys)) # 3. 批量替换 replaced_sentences = [pattern.sub(lambda x: word_list[x.group()], sentence) for sentence in sentences] print(replaced_sentences)
运行后输出和预期完全一致:
["I love learning to code", "coding seems to be interesting using python", "learning how to code using python", "practicing learning how to code is the key"]
原理解释
- 按键长度降序排序:如果存在
a和ab两个键,排序后会先匹配ab,避免a先被替换后ab无法匹配的问题 re.escape处理:如果替换键里包含.、*这类正则特殊字符,也可以正常匹配re.sub回调函数:每匹配到一个键,就直接从替换字典中取出对应值做替换
内容的提问来源于stack exchange,提问作者code_learner
相关产品推荐
相关产品推荐

