为何无法通过正则捕获组提取目标年份数值?
西班牙语文本年份提取问题排查
问题描述
编写Python代码从西班牙语文本中提取年份时,测试示例1无法通过命名捕获组(?P<year>\d*)获取年份数值,使用m1.groups()["\g<year>"]或m1.groups()[0]均返回空值,同时文本替换结果不符合预期。
错误代码
import re input_text_substring = "durante el transcurso del mes de diciembre de 2350" #example 1 #input_text_substring = "durante el transcurso del mes de diciembre del año 2350" #example 2 #input_text_substring = "durante el transcurso del mes 12 2350" #example 3 ##If it is NOT "del año" + "(it doesn't matter how many digits)" or if it is NOT "(it doesn't matter what comes before it)" + "(year of 4 digits)" if not re.search(r"(?:(?:del|de[\s|]*el|el)[\s|]*(?:año|ano)[\s|]*\d*|.*\d{4}$)", input_text_substring): input_text_substring += " de " + datetime.datetime.today().strftime('%Y') + " " #For when no previous phrase indicative of context was indicated, for example "del año" and the number of digits is not 4 some_text = r"(?:(?!\.\s*?\n)[^;])*" #a number of month or some other text without dots . or ;, or \n ((although it must also admit the possible case where there is nothing in the middle or only a whitespace) #we need to capture the group in the position of the last \d* m1 = re.search( r"(?:del[\s|]*mes|de[\s|]*el[\s|]*mes|de[\s|]*mes|\d{2})" + some_text + r"(?P<year>\d*)" , str(input_text_substring), re.IGNORECASE, ) #if m1: identified_year = str(m1.groups()["\g<year>"]) if m1: identified_year = str(m1.groups()[0]) input_text_substring = re.sub( r"(?:del[\s|]*mes|de[\s|]*el[\s|]*mes|de[\s|]*mes|\d{2})" + some_text + r"\d*", identified_year, input_text_substring ) print(repr(identified_year)) print(repr(input_text_substring))
错误输出
'' 'durante el transcurso '
期望输出
'2350' #in example 1, 2 and 3 'durante el transcurso del mes de diciembre 2350' #in example 1 and 2 'durante el transcurso del mes 12 2350' #in example 3
问题原因与修复
1. 正则表达式核心逻辑错误
- 贪婪匹配吃掉年份:
some_text的正则是贪婪模式,会匹配从"mes"之后到文本末尾的所有符合条件的字符,导致后续的(?P<year>\d*)没有剩余字符可匹配,只能得到空字符串。 - 字符类写法错误:
[\s|]会匹配空格或竖线(|在字符类中是普通字符),正确写法应为\s*(匹配任意数量的空格)。 - 命名捕获组访问错误:访问命名捕获组应使用
m1.group("year"),而非m1.groups()["\g<year>"];m1.groups()[0]虽能获取第一个捕获组,但前提是正则正确匹配到内容。
2. 其他代码问题
- 缺少datetime导入:代码中使用
datetime.datetime.today()但未导入datetime模块,会引发NameError。
修复后的代码
import re import datetime input_text_substring = "durante el transcurso del mes de diciembre de 2350" #example 1 #input_text_substring = "durante el transcurso del mes de diciembre del año 2350" #example 2 #input_text_substring = "durante el transcurso del mes 12 2350" #example 3 # 修正正则:去掉错误的[\s|],改为\s*,调整末尾匹配逻辑 if not re.search(r"(?:(?:del|de\s*el|el)\s*(?:año|ano)\s*\d*|.*\d{4}$)", input_text_substring): input_text_substring += " de " + datetime.datetime.today().strftime('%Y') + " " # 改为非贪婪匹配,确保不会吃掉年份;同时限制匹配到年份前的内容 some_text = r"(?:(?!\.\s*?\n|\d{4})[^;])*?" # 修正正则:替换[\s|]为\s*,确保年份捕获组能匹配到数字 m1 = re.search( r"(?:del\s*mes|de\s*el\s*mes|de\s*mes|\d{2})" + some_text + r"(?P<year>\d{4})", input_text_substring, re.IGNORECASE ) if m1: identified_year = m1.group("year") # 正确访问命名捕获组 else: identified_year = "" # 修正替换正则,保留原有的"mes"相关片段,只调整年份部分的格式 input_text_substring = re.sub( r"(del\s*mes|de\s*el\s*mes|de\s*mes|\d{2})" + some_text + r"\d{4}", r"\1" + some_text + r"\g<year>", input_text_substring ) print(repr(identified_year)) print(repr(input_text_substring))
修复说明
- 将
some_text改为非贪婪模式(*?),并添加(?!\d{4})断言,确保匹配到年份前就停止。 - 把所有
[\s|]替换为\s*,正确匹配任意数量的空格。 - 用
m1.group("year")正确访问命名捕获组。 - 修正替换正则,保留原有的"mes"相关片段,只调整年份部分的格式。
内容的提问来源于stack exchange,提问作者Matt095
相关产品推荐
相关产品推荐

