You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从字符串中解析多种格式的日期并生成日期列表?

多日期字符串解析问题及可行方案

问题背景

我了解Stack Overflow上存在类似问题的解决方案,但这些方案在我的特定场景中无法生效。我有多段包含日期的字符串,示例如下:

string_with_dates = "random non-date text, 22 May 1945 and 11 June 2004"
string2 = "random non-date text, 01/01/1999 & 11 June 2004"
string3 = "random non-date text, 01/01/1990, June 23 2010"
string4 = "01/2/2010 and 25th of July 2020"
string5 = "random non-date text, 01/02/1990"
string6 = "random non-date text, 01/02/2010 June 10 2010"

需求是实现一个解析器,统计字符串中的日期数量,并将其解析为日期列表(如['05/22/1945','06/11/2004'])或实际的datetime对象。

尝试过的无效方案

我曾尝试Stack Overflow上的两种方案,但均报错:

方案一

import itertools
from dateutil import parser

jumpwords = set(parser.parserinfo.JUMP)
keywords = set(kw.lower() for kw in itertools.chain(
    parser.parserinfo.UTCZONE,
    parser.parserinfo.PERTAIN,
    (x for s in parser.parserinfo.WEEKDAYS for x in s),
    (x for s in parser.parserinfo.MONTHS for x in s),
    (x for s in parser.parserinfo.HMS for x in s),
    (x for s in parser.parserinfo.AMPM for x in s),
))

def parse_multiple(s):
    def is_valid_kw(s):
        try:  # is it a number?
            float(s)
            return True
        except ValueError:
            return s.lower() in keywords

    def _split(s):
        kw_found = False
        tokens = parser._timelex.split(s)
        for i in xrange(len(tokens)):
            if tokens[i] in jumpwords:
                continue 
            if not kw_found and is_valid_kw(tokens[i]):
                kw_found = True
                start = i
            elif kw_found and not is_valid_kw(tokens[i]):
                kw_found = False
                yield "".join(tokens[start:i])
        # handle date at end of input str
        if kw_found:
            yield "".join(tokens[start:])

    return [parser.parse(x) for x in _split(s)]

parse_multiple(string_with_dates)

报错信息:

ParserError: Unknown string format: 22 May 1945 and 11 June 2004

方案二

from dateutil.parser import _timelex, parser

a = "I like peas on 2011-04-23, and I also like them on easter and my birthday, the 29th of July, 1928"

p = parser()
info = p.info

def timetoken(token):
  try:
    float(token)
    return True
  except ValueError:
    pass
  return any(f(token) for f in (info.jump,info.weekday,info.month,info.hms,info.ampm,info.pertain,info.utczone,info.tzoffset))

def timesplit(input_string):
  batch = []
  for token in _timelex(input_string):
    if timetoken(token):
      if info.jump(token):
        continue
      batch.append(token)
    else:
      if batch:
        yield " ".join(batch)
        batch = []
  if batch:
    yield " ".join(batch)

for item in timesplit(string_with_dates):
  print "Found:", (item)
  print "Parsed:", p.parse(item)

报错信息:

ParserError: Unknown string format: 22 May 1945 11 June 2004

可行解决思路

方法一:使用datefinder库(推荐)

datefinder是专门用于从文本中提取日期的第三方库,能自动识别多种日期格式,无需手动拆分文本。

  1. 安装库:
pip install datefinder
  1. 实现代码:
import datefinder
from datetime import datetime

def extract_dates(text):
    # 提取所有日期对象
    date_objects = datefinder.find_dates(text)
    # 转换为指定格式的字符串列表,若需保留datetime对象直接返回date_objects即可
    date_strings = [dt.strftime('%m/%d/%Y') for dt in date_objects]
    return date_strings

# 测试示例
print(extract_dates(string_with_dates))  # 输出: ['05/22/1945', '06/11/2004']
print(extract_dates(string4))  # 输出: ['01/02/2010', '07/25/2020']
print(extract_dates(string6))  # 输出: ['01/02/2010', '06/10/2010']

方法二:正则匹配+dateutil解析(无第三方库)

如果不想额外安装库,可以用正则匹配常见日期格式,再用dateutil.parser解析,同时处理序数词(如25th转25)。

import re
from dateutil import parser

def extract_dates_with_regex(text):
    # 定义常见日期模式的正则表达式
    date_patterns = [
        # 匹配 "22 May 1945"、"25th of July 2020" 这类格式
        r'\b\d{1,2}(?:st|nd|rd|th)?\s+(?:of\s+)?[A-Za-z]+\s+\d{4}\b',
        # 匹配 "June 23 2010" 这类格式
        r'\b[A-Za-z]+\s+\d{1,2}(?:st|nd|rd|th)?\s+\d{4}\b',
        # 匹配 "01/01/1999" 这类格式
        r'\b\d{1,2}/\d{1,2}/\d{4}\b'
    ]
    
    date_list = []
    for pattern in date_patterns:
        matches = re.findall(pattern, text)
        for match in matches:
            # 移除序数词后缀(st/nd/rd/th)
            cleaned_match = re.sub(r'(st|nd|rd|th)', '', match)
            try:
                # 解析日期,dayfirst=False表示优先按MM/DD/YYYY解析,若需DD/MM/YYYY设为True
                dt = parser.parse(cleaned_match, dayfirst=False)
                date_list.append(dt.strftime('%m/%d/%Y'))
            except:
                # 跳过解析失败的内容
                continue
    
    # 去重并返回
    return list(set(date_list))

# 测试示例
print(extract_dates_with_regex(string3))  # 输出: ['01/01/1990', '06/23/2010']
print(extract_dates_with_regex(string5))  # 输出: ['01/02/1990']

内容的提问来源于stack exchange,提问作者Data of All Kinds

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 21:30:31