如何从字符串中解析多种格式的日期并生成日期列表?
多日期字符串解析问题及可行方案
问题背景
我了解Stack Overflow上存在类似问题的解决方案,但这些方案在我的特定场景中无法生效。我有多段包含日期的字符串,示例如下:
string_with_dates = "random non-date text, 22 May 1945 and 11 June 2004" string2 = "random non-date text, 01/01/1999 & 11 June 2004" string3 = "random non-date text, 01/01/1990, June 23 2010" string4 = "01/2/2010 and 25th of July 2020" string5 = "random non-date text, 01/02/1990" string6 = "random non-date text, 01/02/2010 June 10 2010"
需求是实现一个解析器,统计字符串中的日期数量,并将其解析为日期列表(如['05/22/1945','06/11/2004'])或实际的datetime对象。
尝试过的无效方案
我曾尝试Stack Overflow上的两种方案,但均报错:
方案一
import itertools from dateutil import parser jumpwords = set(parser.parserinfo.JUMP) keywords = set(kw.lower() for kw in itertools.chain( parser.parserinfo.UTCZONE, parser.parserinfo.PERTAIN, (x for s in parser.parserinfo.WEEKDAYS for x in s), (x for s in parser.parserinfo.MONTHS for x in s), (x for s in parser.parserinfo.HMS for x in s), (x for s in parser.parserinfo.AMPM for x in s), )) def parse_multiple(s): def is_valid_kw(s): try: # is it a number? float(s) return True except ValueError: return s.lower() in keywords def _split(s): kw_found = False tokens = parser._timelex.split(s) for i in xrange(len(tokens)): if tokens[i] in jumpwords: continue if not kw_found and is_valid_kw(tokens[i]): kw_found = True start = i elif kw_found and not is_valid_kw(tokens[i]): kw_found = False yield "".join(tokens[start:i]) # handle date at end of input str if kw_found: yield "".join(tokens[start:]) return [parser.parse(x) for x in _split(s)] parse_multiple(string_with_dates)
报错信息:
ParserError: Unknown string format: 22 May 1945 and 11 June 2004
方案二
from dateutil.parser import _timelex, parser a = "I like peas on 2011-04-23, and I also like them on easter and my birthday, the 29th of July, 1928" p = parser() info = p.info def timetoken(token): try: float(token) return True except ValueError: pass return any(f(token) for f in (info.jump,info.weekday,info.month,info.hms,info.ampm,info.pertain,info.utczone,info.tzoffset)) def timesplit(input_string): batch = [] for token in _timelex(input_string): if timetoken(token): if info.jump(token): continue batch.append(token) else: if batch: yield " ".join(batch) batch = [] if batch: yield " ".join(batch) for item in timesplit(string_with_dates): print "Found:", (item) print "Parsed:", p.parse(item)
报错信息:
ParserError: Unknown string format: 22 May 1945 11 June 2004
可行解决思路
方法一:使用datefinder库(推荐)
datefinder是专门用于从文本中提取日期的第三方库,能自动识别多种日期格式,无需手动拆分文本。
- 安装库:
pip install datefinder
- 实现代码:
import datefinder from datetime import datetime def extract_dates(text): # 提取所有日期对象 date_objects = datefinder.find_dates(text) # 转换为指定格式的字符串列表,若需保留datetime对象直接返回date_objects即可 date_strings = [dt.strftime('%m/%d/%Y') for dt in date_objects] return date_strings # 测试示例 print(extract_dates(string_with_dates)) # 输出: ['05/22/1945', '06/11/2004'] print(extract_dates(string4)) # 输出: ['01/02/2010', '07/25/2020'] print(extract_dates(string6)) # 输出: ['01/02/2010', '06/10/2010']
方法二:正则匹配+dateutil解析(无第三方库)
如果不想额外安装库,可以用正则匹配常见日期格式,再用dateutil.parser解析,同时处理序数词(如25th转25)。
import re from dateutil import parser def extract_dates_with_regex(text): # 定义常见日期模式的正则表达式 date_patterns = [ # 匹配 "22 May 1945"、"25th of July 2020" 这类格式 r'\b\d{1,2}(?:st|nd|rd|th)?\s+(?:of\s+)?[A-Za-z]+\s+\d{4}\b', # 匹配 "June 23 2010" 这类格式 r'\b[A-Za-z]+\s+\d{1,2}(?:st|nd|rd|th)?\s+\d{4}\b', # 匹配 "01/01/1999" 这类格式 r'\b\d{1,2}/\d{1,2}/\d{4}\b' ] date_list = [] for pattern in date_patterns: matches = re.findall(pattern, text) for match in matches: # 移除序数词后缀(st/nd/rd/th) cleaned_match = re.sub(r'(st|nd|rd|th)', '', match) try: # 解析日期,dayfirst=False表示优先按MM/DD/YYYY解析,若需DD/MM/YYYY设为True dt = parser.parse(cleaned_match, dayfirst=False) date_list.append(dt.strftime('%m/%d/%Y')) except: # 跳过解析失败的内容 continue # 去重并返回 return list(set(date_list)) # 测试示例 print(extract_dates_with_regex(string3)) # 输出: ['01/01/1990', '06/23/2010'] print(extract_dates_with_regex(string5)) # 输出: ['01/02/1990']
内容的提问来源于stack exchange,提问作者Data of All Kinds
相关产品推荐
相关产品推荐

