You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式在在线测试工具可用但Python中失效,如何修复?

问题:Python正则re.findall返回分组而非完整匹配的问题

我尝试了以下代码:

import re

pattern_time = (
    r"\b\d{1,2}([: .]?)(\d{2})?(\s?)((P|p)?)(.?)((M|m)?)(.?)((next|this)?)(\s?)((tomorrow|today|day|evening|morning|(mid?)night|((after|before)?)(noon))?)\b"
)

test_string = "The meeting is at 3pm today or 5 tomorrow or 7 this afternoon or 00:00 midnight. Let's meet at 11.30 p.M. We can also do 8:45 pm or 1200 hr or 00hr."

matches = re.findall(pattern_time, test_string, re.IGNORECASE)
print(matches)

在Python中得到的输出为:

[('', '', '', 'p', 'p', 'm', '', '', ' ', '', '', '', 'today', 'today', '', '', '', ''), (' ', '', '', '', '', '', '', '', '', '', '', '', 'tomorrow', 'tomorrow', '', '', '', ''), (' ', '', '', '', '', '', '', '', '', 'this', 'this', ' ', 'afternoon', 'afternoon', '', 'after', 'after', 'noon'), (':', '00', ' ', '', '', '', '', '', '', '', '', '', 'midnight', 'midnight', 'mid', '', '', ''), ('.', '30', ' ', 'p', 'p', '.', 'M', 'M', '.', '', '', ' ', '', '', '', '', '', ''), (':', '45', ' ', 'p', 'p', 'm', '', '', ' ', '', '', '', '', '', '', '', '', ''), ('', '00', ' ', '', '', 'h', '', '', 'r', '', '', ' ', '', '', '', '', '', ''), ('', '', '', '', '', 'h', '', '', 'r', '', '', '', '', '', '', '', '', '')]

但在线正则测试工具显示匹配正确,如何在Python中修正此问题?

感谢您的帮助。


解决方案

问题根源是re.findall的特性:当正则表达式包含捕获分组时,它会返回每个分组的匹配结果组成的元组列表,而非完整匹配文本。以下两种方法可以解决:

方法1:改用非捕获分组

将所有不需要单独捕获的分组改为非捕获格式(?:...),这样findall就会返回完整的匹配内容:

import re

pattern_time = (
    r"\b\d{1,2}(?:[: .]?)(?:\d{2})?(?:\s?)(?:(?:P|p)?)(?:.?)(?:(?:M|m)?)(?:.?)(?:(?:next|this)?)(?:\s?)(?:tomorrow|today|day|evening|morning|(?:mid?)night|(?:(?:after|before)?)noon)?\b"
)

test_string = "The meeting is at 3pm today or 5 tomorrow or 7 this afternoon or 00:00 midnight. Let's meet at 11.30 p.M. We can also do 8:45 pm or 1200 hr or 00hr."

matches = re.findall(pattern_time, test_string, re.IGNORECASE)
print(matches)

方法2:使用re.finditer遍历匹配对象

re.finditer返回匹配对象迭代器,通过每个对象的.group()方法可获取完整匹配文本:

import re

pattern_time = (
    r"\b\d{1,2}([: .]?)(\d{2})?(\s?)((P|p)?)(.?)((M|m)?)(.?)((next|this)?)(\s?)((tomorrow|today|day|evening|morning|(mid?)night|((after|before)?)(noon))?)\b"
)

test_string = "The meeting is at 3pm today or 5 tomorrow or 7 this afternoon or 00:00 midnight. Let's meet at 11.30 p.M. We can also do 8:45 pm or 1200 hr or 00hr."

matches = re.finditer(pattern_time, test_string, re.IGNORECASE)
full_matches = [match.group() for match in matches]
print(full_matches)

两种方法最终都会得到预期的完整匹配结果:

['3pm today', '5 tomorrow', '7 this afternoon', '00:00 midnight', '11.30 p.M.', '8:45 pm', '1200 hr', '00hr']

额外优化建议:正则中的(.?)会匹配任意单个字符(包括空),可能导致意外匹配,建议根据需求限制为特定字符,比如改成(?:[ .]?)。

内容的提问来源于stack exchange,提问作者Ursa Major

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 20:33:15