如何获取句子中指定日期列表的起止索引?
问题描述
给定句子:'Foo bar was open on 12.03.2022 and closed on 3.05.22.'
日期列表:list = ['12.03.2022', '4.04.2022', '3.05.22'](注:列表元素需为字符串类型,否则无法直接匹配)
需求:获取列表中能在句子里找到的日期的起止索引,以元组形式返回,期望结果为[(20,29), (45, 51)]。目前已通过正则表达式匹配到日期,但无法获取对应的索引,现有代码如下:
import re DAY = r'(?:(?:0)[1-9]|[12]\d|3[01])' # day can be from 1 to 31 with a leading zero MONTH = r'(?:(?:0)[1-9]|1[0-2])' # month can be 1 to 12 with a leading zero YEAR1 = r'(?:(?:20|)\d{2}|(?:19|){9}[0-9])' # Restricted the year to begin in 20th or 21st century # Also the first two digits may be skipped if data is represented as dd.mm.yy YEAR2 = r'(?:20\d{2}|199[0-9])' BEGIN_LINE1 = r'(?<!\w)' DELIM1 = r'(?:[\,\/\-\._])' DELIM2 = r'(?:[\,\/\-\._])?' # combined, several options NUM_DATE = f"""(?P<date> (?: # DAY MONTH YEAR (?:{BEGIN_LINE1}{DAY}{DELIM1}{MONTH}{DELIM1}{YEAR1}) | (?:{BEGIN_LINE1}{DAY}{DELIM1}{MONTH}) | (?:{BEGIN_LINE1}{MONTH}{DELIM1}{YEAR1}) | (?:{BEGIN_LINE1}{DAY}{DELIM2}{MONTH}{DELIM2}{YEAR2}) | (?:{BEGIN_LINE1}{MONTH}{DELIM2}{YEAR2}) ) )""" myDate = re.compile(f'{NUM_DATE}', re.IGNORECASE | re.VERBOSE | re.UNICODE) def find_date(subject): """_summary_ Args: subject (_type_): _description_ Returns: _type_: _description_ """ if subject is None: return subject dates = list(set(myDate.findall(subject))) return dates
解决方案
要获取匹配日期的起止索引,需用re.finditer()替代findall(),该方法返回匹配对象迭代器,每个匹配对象包含start()(起始索引)和end()(结束索引)方法。
修改后的代码
import re DAY = r'(?:(?:0)[1-9]|[12]\d|3[01])' # 日期:1-31,可带前导零 MONTH = r'(?:(?:0)[1-9]|1[0-2])' # 月份:1-12,可带前导零 YEAR1 = r'(?:(?:20|)\d{2}|(?:19|){9}[0-9])' # 年份限制在20或21世纪,支持简写为yy格式 YEAR2 = r'(?:20\d{2}|199[0-9])' BEGIN_LINE1 = r'(?<!\w)' DELIM1 = r'(?:[\,\/\-\._])' DELIM2 = r'(?:[\,\/\-\._])?' # 组合正则表达式,匹配多种日期格式 NUM_DATE = f"""(?P<date> (?: # 日.月.年 格式 (?:{BEGIN_LINE1}{DAY}{DELIM1}{MONTH}{DELIM1}{YEAR1}) | (?:{BEGIN_LINE1}{DAY}{DELIM1}{MONTH}) | (?:{BEGIN_LINE1}{MONTH}{DELIM1}{YEAR1}) | (?:{BEGIN_LINE1}{DAY}{DELIM2}{MONTH}{DELIM2}{YEAR2}) | (?:{BEGIN_LINE1}{MONTH}{DELIM2}{YEAR2}) ) )""" myDate = re.compile(f'{NUM_DATE}', re.IGNORECASE | re.VERBOSE | re.UNICODE) def find_date_indices(subject, target_dates): if subject is None or not target_dates: return [] # 收集所有匹配的日期及其索引 date_matches = [] for match in myDate.finditer(subject): date_str = match.group('date') start_idx = match.start() end_idx = match.end() date_matches.append((date_str, start_idx, end_idx)) # 筛选出目标列表中存在的日期的索引 result = [] target_set = set(target_dates) for date_str, start, end in date_matches: if date_str in target_set: result.append((start, end)) return result # 测试示例 sentence = 'Foo bar was open on 12.03.2022 and closed on 3.05.22.' target_list = ['12.03.2022', '4.04.2022', '3.05.22'] print(find_date_indices(sentence, target_list)) # 输出: [(20, 29), (45, 51)]
关键说明
re.finditer()的使用:该方法遍历句子中所有匹配的日期,返回的每个Match对象可直接调用start()和end()获取索引。- 目标日期匹配:将目标列表转为集合,可快速判断匹配到的日期是否在目标范围内,提升查找效率。
- 格式一致性:确保目标列表中的日期为字符串类型,避免因类型不匹配导致的匹配失败。
内容的提问来源于stack exchange,提问作者Yana
相关产品推荐
相关产品推荐

