如何用正则表达式匹配并替换单词的单复数形式
实现匹配单词单数、s/es结尾复数的通用替换方案
先回顾基础的替换逻辑:
假设我们有这段语句:
sentence = "A cow runs on the grass"
要把单词cow替换为<SPECIAL>标记,可执行以下代码:
import re to_replace = "cow" # 替换后结果:A <SPECIAL> runs on the grass sentence = re.sub(rf"(?!\B\w)({re.escape(to_replace)})(?<!\w\B)", "<SPECIAL>", sentence, count=1)
如果要同时支持替换加s的复数形式,只需在正则里添加s?:
sentence = "The cows run on the grass" to_replace = "cow" # 替换后结果:The <SPECIAL> run on the grass sentence = re.sub(rf"(?!\B\w)({re.escape(to_replace) + 's?'})(?<!\w\B)", "<SPECIAL>", sentence, count=1)
这段代码能同时匹配单数cow和复数cows。
现在要实现更通用的规则:同时匹配单词的单数形式、以s结尾的复数、以es结尾的复数(忽略其他复杂复数规则),适配以下测试场景:
to_replace = "cow" sentence1 = "The cow runs on the grass" sentence2 = "The cows run on the grass" # -------------- to_replace = "gas" sentence3 = "There are many natural gases"
解决方案
核心是在正则中添加分支化的可选后缀,优先匹配es再匹配s,确保长后缀不会被拆分。完整代码如下:
import re # 测试cow的单数/复数场景 to_replace = "cow" sentence1 = "The cow runs on the grass" sentence2 = "The cows run on the grass" # 构建通用匹配正则 pattern = rf"(?!\B\w)({re.escape(to_replace)})(?:es|s)?(?<!\w\B)" # 替换并输出结果 result1 = re.sub(pattern, "<SPECIAL>", sentence1, count=1) print(result1) # 输出:The <SPECIAL> runs on the grass result2 = re.sub(pattern, "<SPECIAL>", sentence2, count=1) print(result2) # 输出:The <SPECIAL> run on the grass # 测试gas的复数场景 to_replace = "gas" sentence3 = "There are many natural gases" result3 = re.sub(pattern, "<SPECIAL>", sentence3, count=1) print(result3) # 输出:There are many natural <SPECIAL>
正则说明
(?:es|s)?:非捕获组,?:避免捕获后缀内容,|实现分支选择(先匹配es再匹配s),?表示后缀可选(对应单数形式)re.escape(to_replace):自动转义目标单词中的正则特殊字符(如.、*等),避免正则语法错误(?!\B\w)和(?<!\w\B):确保匹配完整单词,不会误替换单词中的子串(比如不会替换coward里的cow)
内容的提问来源于stack exchange,提问作者gorjan
相关产品推荐
相关产品推荐

