You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则表达式高效实现日期时间单数字补0标准化?

标准化自然语言输入的日期时间格式

我需要标准化用户输入的自然语言日期时间数据,将其中单个数字的月、日、时、分补0为两位数字,最终输出格式为YYYY_-_MM_-_DD(hh:mm am or pm)。

我用了Python的re.sub()处理多实例替换,但年、月、日、时、分各有补0/不补0两种情况,总共32种组合,逐个写正则显然不合理,求更优实现方法。

我尝试的代码

import re

input_text = '2000_-_9_-_01 8:1 am'  #example 1
# input_text = '(2000_-_1_-_01) 18:1 pm'  #example 2
# input_text = '(20000_-_12_-_1) (1:1 am)'  #example 3


identificate_hours = r"(?:a\s*las|a\s*la|)\s*(\d{1,2}):(\d{1,2})\s*(?:(am)|(pm)|)"

date_format_00 = r"(\d*)_-_(\d{1,2})_-_(\d{1,2})"
identification_re_0 = r"(?:\(|)\s*" + date_format_00 + r"\s*(?:\)|)\s*(?:a\s*las|a\s*la|)\s*(?:\(|)\s*" + identificate_hours + r"\s*(?:\)|)"

input_text = re.sub(identification_re_0,
                    #lambda m: print(m[2]),
                    lambda m: (f"({m[1]}_-_{m[2]}_-_{m[3]}({m[4] or '00'}:{m[5] or '00'} {m[6] or m[7] or 'am'}))"),
                    input_text, re.IGNORECASE)

print(repr(input_text)) # --> output

期望输出示例

'(2000_-_09_-_01(08:01 am))'   #for example 1
'(2000_-_01_-_01(18:01 pm))'   #for example 2
'(20000_-_12_-_01(01:01 am))'  #for example 3

更优实现方案

不用针对每种补0组合写正则,而是在捕获分组后,通过字符串格式化自动补0。核心是利用Python的格式化语法{:02d}将单个数字转为两位(自动补前导0),两位数字则保持不变。改进后的代码如下:

import re

def standardize_datetime(match):
    # 提取各分组内容
    year = match.group(1)
    month = int(match.group(2))
    day = int(match.group(3))
    hour = int(match.group(4))
    minute = int(match.group(5))
    period = match.group(6) or match.group(7) or 'am'
    period = period.lower()  # 统一为小写格式
    
    # 格式化补0并拼接目标格式
    return f"({year}_-_{month:02d}_-_{day:02d}({hour:02d}:{minute:02d} {period}))"

# 测试三个示例用例
test_cases = [
    '2000_-_9_-_01 8:1 am',
    '(2000_-_1_-_01) 18:1 pm',
    '(20000_-_12_-_1) (1:1 am)'
]

# 正则模式:兼容可选括号、西班牙语时间前缀,捕获核心日期时间字段
pattern = r"(?:\()?\s*(\d*)_-_(\d{1,2})_-_(\d{1,2})\s*(?:\))?\s*(?:a\s*las|a\s*la)?\s*(?:\()?\s*(\d{1,2}):(\d{1,2})\s*(am|pm)?\s*(?:\))?"

for case in test_cases:
    result = re.sub(pattern, standardize_datetime, case, flags=re.IGNORECASE)
    print(repr(result))

代码说明

  1. 正则捕获:通过分组精准捕获年、月、日、时、分及AM/PM字段,同时兼容输入中可选的括号和西班牙语时间前缀a las/a la
  2. 自动补0:将月、日、时、分转为整数后,用{:02d}格式化,自动为单数字补前导0,双数字保持原样
  3. 格式统一:将AM/PM转为小写,确保输出格式一致性
  4. 多实例处理:re.sub()会自动遍历文本,处理所有匹配到的日期时间实例

运行后输出:

'(2000_-_09_-_01(08:01 am))'
'(2000_-_01_-_01(18:01 pm))'
'(20000_-_12_-_01(01:01 am))'

内容的提问来源于stack exchange,提问作者Matt095

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 05:15:20