You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何将自然语言描述的日期转换为标准格式?

解决方案

针对这类包含序数词的自然语言日期(如"The first Wednesday of September, 2021"),可以用Python的datetime和calendar模块做针对性解析,避开datefinder对这类格式的识别盲区,具体实现如下:

  • 先建立序数词到数字的映射(比如first→1、3rd→3)
  • 用正则从文本中提取序数词、星期几、月份、年份这四个核心信息
  • 借助calendar模块生成目标月份的日历,筛选出对应星期几的所有日期,再根据序数词取对应位置的日期

代码示例

import re
import calendar
from datetime import datetime

# 序数词与数字的映射表,覆盖常见写法
ordinal_map = {
    'first': 1, '1st': 1,
    'second': 2, '2nd': 2,
    'third': 3, '3rd': 3,
    'fourth': 4, '4th': 4,
    'fifth': 5, '5th': 5
}

# 匹配目标日期格式的正则,支持大小写不敏感
date_pattern = re.compile(
    r'The (?P<ordinal>\w+|(\d+st|\d+nd|\d+rd|\d+th)) (?P<weekday>\w+) (of|in) (?P<month>\w+), (?P<year>\d{4})',
    re.IGNORECASE
)

def parse_nth_weekday_date(text):
    match = date_pattern.search(text)
    if not match:
        return None
    
    # 提取并标准化各部分信息
    ordinal = match.group('ordinal').lower()
    weekday = match.group('weekday').lower()
    month = match.group('month').lower()
    year = int(match.group('year'))
    
    # 转换序数词为数字
    n = ordinal_map.get(ordinal)
    if not n:
        # 处理数字开头的序数词(如12th)
        n = int(re.match(r'\d+', ordinal).group())
    
    # 映射星期几到calendar模块对应的编号
    weekday_num = {
        'monday': calendar.MONDAY,
        'tuesday': calendar.TUESDAY,
        'wednesday': calendar.WEDNESDAY,
        'thursday': calendar.THURSDAY,
        'friday': calendar.FRIDAY,
        'saturday': calendar.SATURDAY,
        'sunday': calendar.SUNDAY
    }[weekday]
    
    # 映射月份到数字
    month_num = {
        'january': 1, 'february':2, 'march':3, 'april':4,
        'may':5, 'june':6, 'july':7, 'august':8,
        'september':9, 'october':10, 'november':11, 'december':12
    }[month]
    
    # 生成目标月份的日历矩阵,筛选对应星期几的日期
    cal = calendar.monthcalendar(year, month_num)
    target_days = [week[weekday_num] for week in cal if week[weekday_num] != 0]
    
    # 校验序数词有效性(比如部分月份没有第5个星期几)
    if n > len(target_days):
        return None
    target_day = target_days[n-1]
    
    # 格式化为要求的标准格式
    return datetime(year, month_num, target_day).strftime('%B %d, %Y')

# 测试用例
test_texts = [
    "The first Wednesday of September, 2021",
    "The third Monday in July, 2022",
    "The 2nd Tuesday of October, 2023"
]

for text in test_texts:
    print(f"{text} → {parse_nth_weekday_date(text)}")

补充说明

  • 正则表达式可根据实际文本场景调整,比如如果日期描述嵌在长文本中,可以优化匹配逻辑优先提取日期片段
  • 若需要支持更多序数词写法(如twenty-first),只需扩展ordinal_map或增加数字转换逻辑即可

内容的提问来源于stack exchange,提问作者pbthehuman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 14:15:13