You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式分组匹配与非匹配内容的分词实现问题

解决方案

你可以通过两种方式实现需求:

方法一:使用re.split(更简洁)

利用re.split的捕获组特性,将大括号字段作为分隔符并保留它们,之后过滤掉空字符串即可:

import re

s = '{field1}somestring{field2}somestring2{feild3}<somestring3>'
result = [part for part in re.split(r'({[^}]*})', s) if part]
print(result)
# 输出: ['{field1}', 'somestring', '{field2}', 'somestring2', '{feild3}', '<somestring3>']

原理说明:

  • 正则({[^}]*})匹配带大括号的字段,并用括号将其标记为捕获组
  • re.split遇到捕获组时,会将分隔符(也就是大括号字段)也包含在结果列表中
  • 最后用列表推导式过滤掉分割产生的空字符串

方法二:使用re.findall匹配两种模式

编写正则同时匹配大括号字段和非大括号内容,再提取有效结果:

import re

s = '{field1}somestring{field2}somestring2{feild3}<somestring3>'
matches = re.findall(r'({[^}]*})|([^{]+)', s)
result = [item for tup in matches for item in tup if item]
print(result)
# 输出: ['{field1}', 'somestring', '{field2}', 'somestring2', '{feild3}', '<somestring3>']

原理说明:

  • 正则({[^}]*})|([^{]+)分为两个分支:
    1. ({[^}]*})匹配带大括号的字段
    2. ([^{]+)匹配不包含左大括号的连续字符(覆盖所有非大括号内容)
  • re.findall返回的是元组列表,每个元组中只有一个有效元素,因此需要通过列表推导式提取非空元素

为什么你之前的正则不行?

你用的(.*?)({[^}]*})(.*?)存在两个问题:

  1. .*?是惰性匹配,但会匹配空字符串,导致结果出现空值
  2. 该正则只能匹配"非大括号内容+大括号字段+非大括号内容"的结构,无法处理连续的大括号字段或结尾的非大括号内容,因此漏掉了最后的<somestring3>

内容的提问来源于stack exchange,提问作者mike01010

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 01:20:09