You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何构建可互斥提取的多模式正则表达式并映射至对应变量?

如何构建互斥多模式正则表达式以正确提取内容并映射至对应变量?

需求描述

  • 检测并提取指定5种模式(where、what_time、with_whom_or_with_what、why、sense)对应的内容;
  • 提取过程中遇到其他模式的起始标识时立即停止当前提取;
  • 提取结果需分别存入对应模式的变量中。

伪正则模式定义

where模式

(en la|en el|en ) {SIDE} hay que {sense}
hay que {sense} (en la|en el) {SIDE}

what_time模式

(a las|a la) {HOURS} hay que {sense}
hay que {sense} (a las|a la) {HOURS}

with_whom_or_with_what模式

(para ellos|para el|para ellas|para ella|junto a |junto con ) {PERSON} hay que {sense}
hay que {sense} (para) {PERSON}

why模式

(por) {MATTER_OF_WHY} hay que {sense}
hay que {sense} (por) {MATTER_OF_WHY}

输入输出示例

示例1

import re

sense, where, what_time, with_whom_or_with_what, why = "", "", "", "", ""
input_text = "En la montaña de aquel frio lugar hay que estar preparados para largas noches frías junto a tus compañeros"
# 正确输出
sense = "hay que estar preparados para largas noches frías"
where = "En la montaña de aquel frio lugar"
what_time = ""
with_whom_or_with_what = "junto a tus compañeros"
why = ""

示例2

input_text = "Junto a tus compañeros en la montaña de aquel frio lugar hay que estar preparados para largas noches frías"
# 正确输出
sense = "hay que estar preparados para largas noches frías"
where = "en la montaña de aquel frio lugar"
what_time = ""
with_whom_or_with_what = "junto a tus compañeros"
why = ""

示例3

input_text = "Junto a tus compañeros en la montaña de aquel frio lugar hay que estar preparados para largas noches frías porque puede ser peligroso y ocurrir accidentes"
# 正确输出
sense = "hay que estar preparados para largas noches frías"
where = "en la montaña de aquel frio lugar"
what_time = ""
with_whom_or_with_what = "junto a tus compañeros"
why = "porque puede ser peligroso y ocurrir accidentes"

当前问题

现有正则表达式无法正确分组,提取结果混乱,无法映射到对应变量:

import re

sense, where, what_time, with_whom_or_with_what, why = "", "", "", "", ""
input_text = "En la montaña de aquel frio lugar hay que estar preparados para largas noches frías junto a tus compañeros"
regex_pattern = r"((?P<sense>(hay que)\s.+?)|(?P<where>(en la|en el|en)\s.+?)|(?P<what_time>(a las|a la)\s.+?)|(?P<with>(para ellos|para el|para ellas|para ella|junto a|junto con)\s.+?)|(?P<why>(por|porque)\s.+?))(?=\b(en la|en el|en|hay que|a las|a la|para ellos|para el|para ellas|para ella|junto a|junto con|por|porque|$|[!\.\?]))"
n = re.search(regex_pattern, input_text, re.IGNORECASE)
if(n):
    group_list = n.groups()
    print(group_list)

输出结果:

('En la montaña de aquel frio lugar ', None, None, 'En la montaña de aquel frio lugar ', 'En la', None, None, None, None, None, None, 'hay que')

解决方案

单个正则难以处理多模式的位置不确定性和互斥性,建议拆分多个正则分别匹配,先定位核心的sense,再分别匹配前后的其他模式,每个正则通过正向预查确保遇到其他模式时停止提取:

import re

def extract_patterns(input_text):
    sense, where, what_time, with_whom, why = "", "", "", "", ""
    flags = re.IGNORECASE | re.DOTALL
    
    # 提取sense:匹配"hay que"开头,直到遇到其他模式标识或结尾
    sense_match = re.search(r'(?P<sense>hay que .+?(?=\s+(en la|en el|en|a las|a la|para ellos|para el|para ellas|para ella|junto a|junto con|por|porque|$|[!\.\?])))', input_text, flags)
    if sense_match:
        sense = sense_match.group('sense').strip()
    
    # 提取where:匹配sense前后的两种情况
    where_match = re.search(r'(?P<where>(en la|en el|en)\s.+?(?=\s+hay que))', input_text, flags)
    if not where_match:
        where_match = re.search(r'hay que .+?(?P<where>(en la|en el)\s.+?(?=\s+(a las|a la|para ellos|para el|para ellas|para ella|junto a|junto con|por|porque|$|[!\.\?])))', input_text, flags)
    if where_match:
        where = where_match.group('where').strip()
    
    # 提取what_time:匹配sense前后的两种情况
    time_match = re.search(r'(?P<time>(a las|a la)\s.+?(?=\s+hay que))', input_text, flags)
    if not time_match:
        time_match = re.search(r'hay que .+?(?P<time>(a las|a la)\s.+?(?=\s+(en la|en el|en|para ellos|para el|para ellas|para ella|junto a|junto con|por|porque|$|[!\.\?])))', input_text, flags)
    if time_match:
        what_time = time_match.group('time').strip()
    
    # 提取with_whom_or_with_what:匹配sense前后的两种情况
    with_match = re.search(r'(?P<with>(para ellos|para el|para ellas|para ella|junto a |junto con )\s.+?(?=\s+hay que))', input_text, flags)
    if not with_match:
        with_match = re.search(r'hay que .+?(?P<with>(para)\s.+?(?=\s+(en la|en el|en|a las|a la|por|porque|$|[!\.\?])))', input_text, flags)
    if with_match:
        with_whom = with_match.group('with').strip()
    
    # 提取why:匹配sense前后的两种情况
    why_match = re.search(r'(?P<why>(por)\s.+?(?=\s+hay que))', input_text, flags)
    if not why_match:
        why_match = re.search(r'hay que .+?(?P<why>(por|porque)\s.+?(?=$|[!\.\?]))', input_text, flags)
    if why_match:
        why = why_match.group('why').strip()
    
    return sense, where, what_time, with_whom, why

# 测试示例1
input_text1 = "En la montaña de aquel frio lugar hay que estar preparados para largas noches frías junto a tus compañeros"
sense1, where1, time1, with1, why1 = extract_patterns(input_text1)
print("示例1结果:")
print(f"sense: '{sense1}'")
print(f"where: '{where1}'")
print(f"what_time: '{time1}'")
print(f"with_whom_or_with_what: '{with1}'")
print(f"why: '{why1}'")

# 测试示例2
input_text2 = "Junto a tus compañeros en la montaña de aquel frio lugar hay que estar preparados para largas noches frías"
sense2, where2, time2, with2, why2 = extract_patterns(input_text2)
print("\n示例2结果:")
print(f"sense: '{sense2}'")
print(f"where: '{where2}'")
print(f"what_time: '{time2}'")
print(f"with_whom_or_with_what: '{with2}'")
print(f"why: '{why2}'")

# 测试示例3
input_text3 = "Junto a tus compañeros en la montaña de aquel frio lugar hay que estar preparados para largas noches frías porque puede ser peligroso y ocurrir accidentes"
sense3, where3, time3, with3, why3 = extract_patterns(input_text3)
print("\n示例3结果:")
print(f"sense: '{sense3}'")
print(f"where: '{where3}'")
print(f"what_time: '{time3}'")
print(f"with_whom_or_with_what: '{with3}'")
print(f"why: '{why3}'")

代码说明

  1. 核心思路:先提取sense作为锚点,再分别匹配其他模式在sense前后的两种位置情况;
  2. 正向预查:每个正则使用(?=...)确保提取到其他模式的起始标识时立即停止;
  3. 大小写不敏感:通过re.IGNORECASE处理输入文本的大小写差异;
  4. 结果处理:每个匹配结果做strip()去除多余空格,保证输出整洁。

运行上述代码后,三个示例均可得到符合预期的输出结果。

内容的提问来源于stack exchange,提问作者user18051870

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 18:31:25