You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何无法通过正则捕获组提取目标年份数值?

西班牙语文本年份提取问题排查

问题描述

编写Python代码从西班牙语文本中提取年份时,测试示例1无法通过命名捕获组(?P<year>\d*)获取年份数值,使用m1.groups()["\g<year>"]或m1.groups()[0]均返回空值,同时文本替换结果不符合预期。

错误代码

import re

input_text_substring = "durante el transcurso del mes de diciembre de 2350" #example 1
#input_text_substring = "durante el transcurso del mes de diciembre del año 2350" #example 2
#input_text_substring = "durante el transcurso del mes 12 2350" #example 3

##If it is NOT "del año" + "(it doesn't matter how many digits)" or if it is NOT "(it doesn't matter what comes before it)" + "(year of 4 digits)"
if not re.search(r"(?:(?:del|de[\s|]*el|el)[\s|]*(?:año|ano)[\s|]*\d*|.*\d{4}$)", input_text_substring):
    input_text_substring += " de " + datetime.datetime.today().strftime('%Y') + " "

#For when no previous phrase indicative of context was indicated, for example "del año" and the number of digits is not 4

some_text = r"(?:(?!\.\s*?\n)[^;])*" #a number of month or some other text without dots .  or ;, or \n ((although it must also admit the possible case where there is nothing in the middle or only a whitespace)

#we need to capture the group in the position of the last \d*
m1 = re.search( r"(?:del[\s|]*mes|de[\s|]*el[\s|]*mes|de[\s|]*mes|\d{2})" + some_text + r"(?P<year>\d*)" , str(input_text_substring), re.IGNORECASE, )
#if m1: identified_year = str(m1.groups()["\g<year>"])
if m1: identified_year = str(m1.groups()[0])

input_text_substring = re.sub( r"(?:del[\s|]*mes|de[\s|]*el[\s|]*mes|de[\s|]*mes|\d{2})" + some_text + r"\d*", identified_year, input_text_substring )


print(repr(identified_year))
print(repr(input_text_substring))

错误输出

''
'durante el transcurso '

期望输出

'2350' #in example 1, 2 and 3
'durante el transcurso del mes de diciembre 2350' #in example 1 and 2
'durante el transcurso del mes 12 2350' #in example 3

问题原因与修复

1. 正则表达式核心逻辑错误

  • 贪婪匹配吃掉年份:some_text的正则是贪婪模式,会匹配从"mes"之后到文本末尾的所有符合条件的字符,导致后续的(?P<year>\d*)没有剩余字符可匹配,只能得到空字符串。
  • 字符类写法错误:[\s|]会匹配空格或竖线(|在字符类中是普通字符),正确写法应为\s*(匹配任意数量的空格)。
  • 命名捕获组访问错误:访问命名捕获组应使用m1.group("year"),而非m1.groups()["\g<year>"];m1.groups()[0]虽能获取第一个捕获组,但前提是正则正确匹配到内容。

2. 其他代码问题

  • 缺少datetime导入:代码中使用datetime.datetime.today()但未导入datetime模块,会引发NameError。

修复后的代码

import re
import datetime

input_text_substring = "durante el transcurso del mes de diciembre de 2350" #example 1
#input_text_substring = "durante el transcurso del mes de diciembre del año 2350" #example 2
#input_text_substring = "durante el transcurso del mes 12 2350" #example 3

# 修正正则:去掉错误的[\s|],改为\s*,调整末尾匹配逻辑
if not re.search(r"(?:(?:del|de\s*el|el)\s*(?:año|ano)\s*\d*|.*\d{4}$)", input_text_substring):
    input_text_substring += " de " + datetime.datetime.today().strftime('%Y') + " "

# 改为非贪婪匹配,确保不会吃掉年份;同时限制匹配到年份前的内容
some_text = r"(?:(?!\.\s*?\n|\d{4})[^;])*?"

# 修正正则:替换[\s|]为\s*,确保年份捕获组能匹配到数字
m1 = re.search(
    r"(?:del\s*mes|de\s*el\s*mes|de\s*mes|\d{2})" + some_text + r"(?P<year>\d{4})",
    input_text_substring,
    re.IGNORECASE
)

if m1:
    identified_year = m1.group("year")  # 正确访问命名捕获组
else:
    identified_year = ""

# 修正替换正则,保留原有的"mes"相关片段,只调整年份部分的格式
input_text_substring = re.sub(
    r"(del\s*mes|de\s*el\s*mes|de\s*mes|\d{2})" + some_text + r"\d{4}",
    r"\1" + some_text + r"\g<year>",
    input_text_substring
)

print(repr(identified_year))
print(repr(input_text_substring))

修复说明

  • 将some_text改为非贪婪模式(*?),并添加(?!\d{4})断言,确保匹配到年份前就停止。
  • 把所有[\s|]替换为\s*,正确匹配任意数量的空格。
  • 用m1.group("year")正确访问命名捕获组。
  • 修正替换正则,保留原有的"mes"相关片段,只调整年份部分的格式。

内容的提问来源于stack exchange,提问作者Matt095

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 15:20:26