You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则表达式匹配特定姓名格式并替换的问题求助

Python正则表达式匹配特定姓名格式并替换的问题求助

嗨,我来帮你搞定这个问题!你的需求是匹配像"Jane and John Smith"这种「名 + and + 名 + 姓」的格式,同时排除"Jane Smith and John Smith"这种「完整姓名 + and + 完整姓名」的情况,对吧?

你现在写的正则r'([A-Z][a-z]+)\s+and\s+([A-Z][a-z]+)\s+([A-Z][a-z]+)'之所以会误匹配,核心问题是没限制and前面的名字不能有前置的大写开头单词——比如在"Jane Smith and John Smith"里,Smith是大写开头的,你的正则会错误捕获到Smith and John Smith这种不符合需求的片段。

解决办法很简单,给第一个名字分组加上负向零宽断言,确保它的前面没有另一个大写开头的单词(也就是避免前面是完整姓名的情况)。修改后的正则如下:

r'(?<!\b[A-Z][a-z]+\s)([A-Z][a-z]+)\s+and\s+([A-Z][a-z]+)\s+([A-Z][a-z]+)'

这里的(?<!\b[A-Z][a-z]+\s)是关键:它表示当前位置的前面,不能是「大写开头+小写字母的单词+空格」,这样就完美排除了前面已经有一个名字的场景,精准锁定你要的格式。

替换的时候,只需要把匹配到的内容转换成「分组1 分组3 and 分组2 分组3」就行,在Python里用re.sub()实现的话,替换字符串写r'\1 \3 and \2 \3'就可以。

给你个完整的示例代码参考:

import re

text = """
I met Jane and John Smith yesterday, but Jane Smith and John Smith were not there.
Also, Alice and Bob Brown came to the party, but Alice Brown and Bob Brown left early.
"""

pattern = r'(?<!\b[A-Z][a-z]+\s)([A-Z][a-z]+)\s+and\s+([A-Z][a-z]+)\s+([A-Z][a-z]+)'
result = re.sub(pattern, r'\1 \3 and \2 \3', text)

print(result)

运行后输出的结果会是:

I met Jane Smith and John Smith yesterday, but Jane Smith and John Smith were not there.
Also, Alice Brown and Bob Brown came to the party, but Alice Brown and Bob Brown left early.

这样就精准替换了你想要的格式,同时完全不会触动已经是完整姓名的部分!如果之后遇到名字带连字符、姓是多单词的特殊情况,再根据实际需求微调正则就好,当前方案完全匹配你给出的[A-Z][a-z]+格式要求。

备注:内容来源于stack exchange,提问作者Bahar S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.21 11:55:27