You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中高效解析多结构文件并匹配成对大括号

a lot of objects node

object triplex_meter {
name R2-12-47-3_tm_403;
phases AS;
voltage_1 120;
voltage_2 120;
voltage_N 0;
nominal_voltage 120;
}
....

a lot of object triplex_meter

object triplex_line {
groupid Triplex_Line;
name R2-12-47-3_tl_409;
phases AS;
from R2-12-47-3_tn_409;
to R2-12-47-3_tm_409;
length 30;
configuration triplex_line_configuration_1;
}
...

a lot of object triplex_meter

#some nested objects...awh...

我希望能快速匹配`{`和`}`,聚焦处理内部的对象类型,解析后实现如下逻辑:
```python
if obj_type == "node":
        # to do 1
elif obj_type == "triplex_meter":
        # to do 2 

这种结构看起来不难,但我不确定具体该从哪着手。


回答

嘿,这个问题我之前处理类似的配置文件时碰到过,给你几个实用的思路,分情况选就行:

方法一:基于栈的大括号匹配(最通用,支持嵌套)

这是处理括号匹配的经典方案,不管有没有嵌套都能稳得住。核心思路就是用栈追踪括号的深度:

  • 逐行扫描文件,碰到object [类型] {这种起始行时,记录当前对象类型,同时把栈深度+1
  • 之后每遇到一个{就加栈深度,遇到}就减深度
  • 当栈深度回到0的时候,说明当前对象的内容已经全部读取完毕,这时候就可以根据记录的类型执行对应的处理逻辑

给你写了个Python的示例代码,你可以直接改改就能用:

def parse_objects(file_path):
    current_obj_type = None
    brace_stack = []
    current_obj_content = []
    
    with open(file_path, 'r') as f:
        for line in f:
            stripped_line = line.strip()
            # 跳过注释和空行
            if stripped_line.startswith('#') or not stripped_line:
                continue
            
            # 匹配对象起始行,提取类型
            if stripped_line.startswith('object'):
                parts = stripped_line.split()
                if len(parts) >= 2 and parts[-1] == '{':
                    current_obj_type = parts[1]
                    brace_stack.append('{')
                    current_obj_content.append(line)
                    continue
            
            # 处理行内的括号
            if '{' in stripped_line:
                brace_stack.extend(['{' for _ in stripped_line.split('{') if _])
            if '}' in stripped_line:
                for _ in stripped_line.split('}'):
                    if brace_stack:
                        brace_stack.pop()
                        # 栈空=当前对象结束
                        if not brace_stack:
                            current_obj_content.append(line)
                            # 执行你的类型分支逻辑
                            if current_obj_type == 'node':
                                print(f"Processing node content:\n{''.join(current_obj_content)}")
                                # 这里替换成你的to do 1逻辑
                            elif current_obj_type == 'triplex_meter':
                                print(f"Processing triplex_meter content:\n{''.join(current_obj_content)}")
                                # 这里替换成你的to do 2逻辑
                            elif current_obj_type == 'triplex_line':
                                print(f"Processing triplex_line content:\n{''.join(current_obj_content)}")
                                # 这里替换成对应的处理逻辑
                            # 重置状态,准备下一个对象
                            current_obj_type = None
                            current_obj_content = []
                        else:
                            current_obj_content.append(line)
                continue
            
            # 如果当前处于对象内部,收集内容行
            if current_obj_type is not None:
                current_obj_content.append(line)

方法二:正则表达式(适合无嵌套的简单场景)

如果你的文件里没有嵌套的大括号(看你提到有嵌套,这个方法就不太适用,但如果大部分对象都是平级的可以试试),可以用正则一次性匹配整个对象块,代码会更简洁:

import re

def parse_with_regex(file_path):
    with open(file_path, 'r') as f:
        content = f.read()
    
    # 匹配所有object块,注意:嵌套大括号会导致匹配失败
    pattern = re.compile(r'object (\w+) \{([\s\S]+?)\}', re.MULTILINE)
    matches = pattern.findall(content)
    
    for obj_type, obj_content in matches:
        # 清理内容:去掉注释和空行
        cleaned_content = '\n'.join([
            line.strip() for line in obj_content.split('\n') 
            if line.strip() and not line.strip().startswith('#')
        ])
        if obj_type == 'node':
            print(f"Processing node:\n{cleaned_content}")
            # 替换成你的to do 1逻辑
        elif obj_type == 'triplex_meter':
            print(f"Processing triplex_meter:\n{cleaned_content}")
            # 替换成你的to do 2逻辑
        elif obj_type == 'triplex_line':
            print(f"Processing triplex_line:\n{cleaned_content}")
            # 替换成对应的处理逻辑

⚠️ 注意:这个正则用了非贪婪匹配[\s\S]+?,但如果对象内部有嵌套的{},比如某个属性值里带大括号,正则会提前匹配到第一个}就结束,导致内容截断,所以嵌套场景还是优先用栈的方法。

额外优化建议

  • 如果文件特别大,一定要用逐行读取的栈方法,别一次性读入整个文件,避免内存占用过高
  • 可以把对象内容进一步解析成字典,比如把name R2-12-47-3_node_453;转成{'name': 'R2-12-47-3_node_453'},后续处理属性会更方便:
def parse_obj_content(content_lines):
    obj_dict = {}
    for line in content_lines:
        stripped = line.strip()
        if not stripped or stripped.startswith('#') or '{' in stripped or '}' in stripped:
            continue
        # 分割属性名和值,支持值带空格的情况(用maxsplit=1)
        parts = stripped.rstrip(';').split(maxsplit=1)
        if len(parts) == 2:
            key, value = parts
            obj_dict[key] = value
    return obj_dict

然后在栈方法里,当对象结束时调用这个函数,就能拿到结构化的字典,处理逻辑会清晰很多。

内容的提问来源于stack exchange,提问作者Xu Siyuan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 17:01:49