You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python最短正则匹配(.*?)未达预期结果,求技术指导

问题分析与解决办法

我一眼就看出问题出在哪了——你的正则表达式虽然用了非贪婪匹配.*?,但它还是会从第一个<TIMEX3开始匹配,一直延伸到包含24 hours的那个</TIMEX3>标签结束,把中间一大段内容全给覆盖了!

为什么会这样?

.*?是非贪婪匹配没错,但它只是尽可能少地匹配字符,可它的匹配范围并没有被限制在单个<TIMEX3>标签内。在你的文本里,第一个<TIMEX3>标签之后,直到第二个</TIMEX3>之前的所有内容,都会被.*?匹配到,所以整个被替换的部分其实是:

<TIMEX3 tid="t4" type="DATE" value="2013-03-21">the 90 years</TIMEX3> since Rebecca Felton of Georgia became the first woman in the United States Senate - sworn in for a mere <TIMEX3 tid="t5" type="DATE" value="2013-03-21">24 hours</TIMEX3>

替换成24 hours后,自然就只剩下In 24 hours - women remain an anomaly in the upper chamber.了。

正确的正则写法

要解决这个问题,你需要限制<TIMEX3>标签的匹配范围,确保只匹配到包含24 hours的那个标签。把正则里的.*?换成[^>]*即可——[^>]*表示匹配除了>`之外的任意字符,这样就不会跨标签匹配了:

import re

text = 'In <TIMEX3 tid="t4" type="DATE" value="2013-03-21">the 90 years</TIMEX3> since Rebecca Felton of Georgia became the first woman in the United States Senate - sworn in for a mere <TIMEX3 tid="t5" type="DATE" value="2013-03-21">24 hours</TIMEX3> - women remain an anomaly in the upper chamber.'
result = re.sub(r"<TIMEX3[^>]*>24 hours</TIMEX3>", "24 hours", text)
print(result)

输出结果

In <TIMEX3 tid="t4" type="DATE" value="2013-03-21">the 90 years</TIMEX3> since Rebecca Felton of Georgia became the first woman in the United States Senate - sworn in for a mere 24 hours - women remain an anomaly in the upper chamber.

完美符合你的预期!

内容的提问来源于stack exchange,提问作者Barun Patra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 11:43:13