You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3正则提取NAME_分组信息问题求助

Python正则按NAME_ ID分组提取内容的解决方法

需求说明

需要从以下文本中,按NAME_后的ID分组提取内容,每个分组包含对应ID的NAME_行及所有后续同ID的EX_行:

AB_ NAME_ 111 "fruit";;
AB_ EX_ 111 first_fruit "banana";;
AB_ EX_ 111 second_fruit_info "Do you like 
apple

or grape?";;
AB_ EX_ 111 third_fruit "tomato";;
AB_ NAME_ 120 "food";;
AB_ NAME_ 130 "clothes";;
AB_ EX_ 130 first_clothes "t-shirt";; 

期望得到3个分组,分别对应ID 111、120、130的相关内容。

原正则的问题

你使用的正则AB_ NAME_ \d+ .+;\n(AB_ EX_ \d+ (.|\n)+;\n)*存在以下问题:

  • 未关联NAME_和EX_的ID,可能匹配到其他ID的EX_行
  • 贪婪匹配(.+、(.|\n)+)会过度捕获,甚至包含后续的NAME_行
  • 换行符处理逻辑不精准,无法正确匹配跨行的EX_内容

正确的正则方案

使用带反向引用的非贪婪匹配,并结合re.DOTALL标志(让.匹配换行符),正则表达式如下:

AB_ NAME_ (\d+) .+?;;(?:\nAB_ EX_ \1 .+?;;)*

正则解析

  • AB_ NAME_ (\d+) .+?;;:匹配NAME_行,捕获ID到分组1,.+?非贪婪匹配直到该行末尾的;;
  • (?:\nAB_ EX_ \1 .+?;;)*:非捕获组,匹配零个或多个同ID的EX_行,\1引用前面捕获的ID,确保EX_行与NAME_行ID一致,.+?;;处理跨行内容

Python代码实现

import re

text = '''AB_ NAME_ 111 "fruit";;
AB_ EX_ 111 first_fruit "banana";;
AB_ EX_ 111 second_fruit_info "Do you like 
apple

or grape?";;
AB_ EX_ 111 third_fruit "tomato";;
AB_ NAME_ 120 "food";;
AB_ NAME_ 130 "clothes";;
AB_ EX_ 130 first_clothes "t-shirt";; '''

# 编译正则,启用DOTALL模式处理跨行内容
pattern = re.compile(r'AB_ NAME_ (\d+) .+?;;(?:\nAB_ EX_ \1 .+?;;)*', re.DOTALL)
# 查找所有匹配的分组
groups = [match.group(0) for match in pattern.finditer(text)]

# 输出结果
for idx, content in enumerate(groups, 1):
    print(f"{idx})")
    print("<pre><code>")
    print(content)
    print("</code></pre>")

运行后会输出你期望的3个分组内容。

内容的提问来源于stack exchange,提问作者Stella

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 21:01:54