Python最短正则匹配(.*?)未达预期结果,求技术指导
我一眼就看出问题出在哪了——你的正则表达式虽然用了非贪婪匹配.*?,但它还是会从第一个<TIMEX3开始匹配,一直延伸到包含24 hours的那个</TIMEX3>标签结束,把中间一大段内容全给覆盖了!
为什么会这样?
.*?是非贪婪匹配没错,但它只是尽可能少地匹配字符,可它的匹配范围并没有被限制在单个<TIMEX3>标签内。在你的文本里,第一个<TIMEX3>标签之后,直到第二个</TIMEX3>之前的所有内容,都会被.*?匹配到,所以整个被替换的部分其实是:
<TIMEX3 tid="t4" type="DATE" value="2013-03-21">the 90 years</TIMEX3> since Rebecca Felton of Georgia became the first woman in the United States Senate - sworn in for a mere <TIMEX3 tid="t5" type="DATE" value="2013-03-21">24 hours</TIMEX3>
替换成24 hours后,自然就只剩下In 24 hours - women remain an anomaly in the upper chamber.了。
正确的正则写法
要解决这个问题,你需要限制<TIMEX3>标签的匹配范围,确保只匹配到包含24 hours的那个标签。把正则里的.*?换成[^>]*即可——[^>]*表示匹配除了>`之外的任意字符,这样就不会跨标签匹配了:
import re text = 'In <TIMEX3 tid="t4" type="DATE" value="2013-03-21">the 90 years</TIMEX3> since Rebecca Felton of Georgia became the first woman in the United States Senate - sworn in for a mere <TIMEX3 tid="t5" type="DATE" value="2013-03-21">24 hours</TIMEX3> - women remain an anomaly in the upper chamber.' result = re.sub(r"<TIMEX3[^>]*>24 hours</TIMEX3>", "24 hours", text) print(result)
输出结果
In <TIMEX3 tid="t4" type="DATE" value="2013-03-21">the 90 years</TIMEX3> since Rebecca Felton of Georgia became the first woman in the United States Senate - sworn in for a mere 24 hours - women remain an anomaly in the upper chamber.
完美符合你的预期!
内容的提问来源于stack exchange,提问作者Barun Patra

