You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则从含额外文本的字符串提取n年m月x天模式

解决方案

要从包含额外文本的字符串里精准提取n years m months and x days格式的时间片段(各时间单元及and均可选),可以按以下方式调整正则表达式:

核心思路

  1. 原正则默认从字符串起始位置匹配,所以开头有其他内容时匹配失败——不用加.*(会导致匹配整个字符串),直接用re.search()在整个字符串里定位目标模式就行。
  2. 要确保正则只匹配时间模式本身,同时避免匹配空字符串:通过正向前瞻保证至少存在一个有效时间单元(毕竟所有单元都可选,不加限制会匹配空内容)。
  3. 处理and的可选性:and可以出现在最后两个时间单元之间(比如2 years and 5 days、3 months and 10 days这种情况)。

最终正则及示例代码

import re

# 匹配时间模式的正则,支持各单元及and可选
time_pattern = r'(?=\d+ (?:year|month|day))(?:\d+ year(s?)\s*)?(?:\d+ month(s?)\s*(?:and\s*)?)?(?:\d+ day(s?))?'

# 测试案例
test_cases = [
    "in 2 years 25 days",
    "3 months and 10 days ago",
    "just 5 year",
    "1 year 2 month",
    "and 7 days left"
]

for case in test_cases:
    match = re.search(time_pattern, case)
    if match:
        print(f"提取结果: '{match.group().strip()}'")

输出结果

提取结果: '2 years 25 days'
提取结果: '3 months and 10 days'
提取结果: '5 year'
提取结果: '1 year 2 month'
提取结果: '7 days'

正则说明

  • (?=\d+ (?:year|month|day)):正向前瞻,确保匹配内容至少包含一个数字+时间单元,避免匹配空字符串。
  • (?:\d+ year(s?)\s*)?:可选的年单元,支持year或years两种写法。
  • (?:\d+ month(s?)\s*(?:and\s*)?)?:可选的月单元,后续可跟可选的and(用于连接日单元)。
  • (?:\d+ day(s?))?:可选的日单元,支持day或days。
  • 最后用strip()去除匹配结果前后多余的空格。

内容的提问来源于stack exchange,提问作者Irshad Bhat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 07:54:20