You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则提取through与at间日期 无匹配时返回完整原字符串

Python正则兼容两种场景提取日期方案

规则说明

待处理字符串共两类,提取规则如下:

  • 字符串同时包含through和at关键词:提取through之后、at之前的日期部分
  • 字符串本身为纯日期格式(无上述两个关键词):直接返回完整原字符串

测试样本:

Beginning through June 18, 2022 at Noon standard time
Jan 20, 2022
Beginning through April 26, 2022 at 12:01 a.m. standard time

预期输出:

June 18, 2022
Jan 20, 2022
April 26, 2022

原有方案缺陷

原有正则仅能覆盖带关键词的长文本场景:

s ="Beginning through June 18, 2022 at Noon standard time"
re.search(r'(.*through)(.*) (at.*)', s).group(2)

传入纯日期字符串(如"June 18, 2022")时,会因为匹配不到through/at关键词返回空值,无法拿到正确结果。

实现代码

通过正则分支逻辑同时覆盖两种匹配场景:优先匹配两个关键词中间的日期内容,匹配不到时校验是否为纯日期格式直接返回,代码如下:

import re

def extract_date(input_str):
    # 正则包含两个匹配分支:
    # 1. 非贪婪匹配through和at之间的内容,避免捕获多余字符
    # 2. 匹配开头到结尾的纯日期格式(月份+日期+逗号+年份)
    match_res = re.search(
        r"through\s+(.*?)\s+at|^([A-Za-z]+\s+\d{1,2},\s+\d{4})$",
        input_str.strip()
    )
    if not match_res:
        return ""
    # 返回非空分组的结果,自动去除前后空白字符
    return (match_res.group(1) or match_res.group(2)).strip()

# 测试验证
if __name__ == "__main__":
    test_list = [
        "Beginning through June 18, 2022 at Noon standard time",
        "Jan 20, 2022",
        "Beginning through April 26, 2022 at 12:01 a.m. standard time"
    ]
    for item in test_list:
        print(extract_date(item))

运行输出和预期完全一致:

June 18, 2022
Jan 20, 2022
April 26, 2022

提示:如果后续需要适配其他格式的纯日期(比如月份全称、带星期前缀的日期),只需要调整第二个分支的日期匹配规则即可。

内容的提问来源于stack exchange,提问作者g_p

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 08:21:18