使用Python str.replace替换单复数相关字符串时如何避免拼写错误
问题原因
直接调用Python内置的
str.replace()方法会执行全局模糊子串匹配,没有校验匹配内容是否为独立单词。你要替换的短关键词it是长关键词its的前缀,所以替换it时会命中its的前两个字符,替换后剩下末尾的s,最终就出现了diabetess这类多余后缀的错误。
另外原代码还存在几个冗余/错误点:
- 遍历长度时误用了未定义的变量
repl1- 多处重复append同一个
xxx变量- 无意义的numpy数组转换操作
解决方案
使用正则表达式的**单词边界锚点\b**实现独立单词匹配,只有完全匹配独立单词时才执行替换,避免命中其他长单词的子串。
修正后代码
import re from collections import OrderedDict def replacing(): texter = [] repl = ['diabetes', 'mellitus', 'dm'] txt = "tell me if its can also cause coronavirus" # 要替换的独立目标词列表 target_words = ["its", "it", "them", "the same", "this"] for p in repl: for word in target_words: # 用\b包裹目标词匹配独立单词,re.escape避免特殊字符干扰 replaced = re.sub(rf'\b{re.escape(word)}\b', p, txt) texter.append(replaced) mm = list(OrderedDict.fromkeys(texter)) print(mm) replacing()
运行结果
去重后输出为:
['tell me if diabetes can also cause coronavirus', 'tell me if mellitus can also cause coronavirus', 'tell me if dm can also cause coronavirus']
完全消除了多余后缀的拼写错误。
内容的提问来源于stack exchange,提问作者lobjc
相关产品推荐
相关产品推荐

