You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中使用正则匹配平衡嵌套{}提取复杂字符串内容的方案

问题根因

你整合正则时存在两个核心错误:

  • (?R)的作用是递归整个完整正则表达式,单独匹配括号时整个正则就是括号规则所以生效,整合后整个正则包含extend前缀规则,递归时会再次要求匹配extend,自然无法正确匹配嵌套括号。
  • 整合后的正则多了冗余规则extend\s+([^n]*),会错误吞掉extend后面的phase字段,导致后续匹配逻辑失效。

解决方案

改用regex库的DEFINE块封装括号平衡的子规则,递归时仅调用局部子规则即可,修正后代码如下:

  1. 先安装第三方正则库:pip install regex
  2. 运行代码:
import regex as re

my_string = """
extend mineral Uraninite {
    kinetics {
        rate = -3.2e-08 mol/m2/s
        area = Uraninite
        y-term, species = Uraninite
        w-term {
            species = H[+]
            power = 0.37
        }
    }
    kinetics {
        rate = 3.2e-09 mol/m2/s
        area = Uraninite
        y-term, species = Uraninite
        w-term {
            species = H[+]
            power = 0.37
        }
    }
}
"""

pattern = re.compile(
    r"extend\s+"
    r"(?:(?P<phase>colloid|mineral|basis|isotope|solid-solution)\s+)?"
    r"(?P<species>[^\n ]+)\s+"
    r"{(?P<content>(?&brace_group))}"
    # 定义括号平衡的独立子规则,仅递归该部分
    r"(?(DEFINE)(?P<brace_group>(?:[^{}]++|(?&brace_group))*+))"
)

extend_list = [m.groupdict() for m in pattern.finditer(my_string)]
# 测试输出
for item in extend_list:
    print("phase:", item["phase"])
    print("species:", item["species"])
    print("content:\n", item["content"])

内容的提问来源于stack exchange,提问作者Antoine Collet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 20:36:07