如何在Python中高效解析多结构文件并匹配成对大括号
object triplex_meter {
name R2-12-47-3_tm_403;
phases AS;
voltage_1 120;
voltage_2 120;
voltage_N 0;
nominal_voltage 120;
}
....
object triplex_line {
groupid Triplex_Line;
name R2-12-47-3_tl_409;
phases AS;
from R2-12-47-3_tn_409;
to R2-12-47-3_tm_409;
length 30;
configuration triplex_line_configuration_1;
}
...
#some nested objects...awh...
我希望能快速匹配`{`和`}`,聚焦处理内部的对象类型,解析后实现如下逻辑: ```python if obj_type == "node": # to do 1 elif obj_type == "triplex_meter": # to do 2
这种结构看起来不难,但我不确定具体该从哪着手。
回答
嘿,这个问题我之前处理类似的配置文件时碰到过,给你几个实用的思路,分情况选就行:
方法一:基于栈的大括号匹配(最通用,支持嵌套)
这是处理括号匹配的经典方案,不管有没有嵌套都能稳得住。核心思路就是用栈追踪括号的深度:
- 逐行扫描文件,碰到
object [类型] {这种起始行时,记录当前对象类型,同时把栈深度+1 - 之后每遇到一个
{就加栈深度,遇到}就减深度 - 当栈深度回到0的时候,说明当前对象的内容已经全部读取完毕,这时候就可以根据记录的类型执行对应的处理逻辑
给你写了个Python的示例代码,你可以直接改改就能用:
def parse_objects(file_path): current_obj_type = None brace_stack = [] current_obj_content = [] with open(file_path, 'r') as f: for line in f: stripped_line = line.strip() # 跳过注释和空行 if stripped_line.startswith('#') or not stripped_line: continue # 匹配对象起始行,提取类型 if stripped_line.startswith('object'): parts = stripped_line.split() if len(parts) >= 2 and parts[-1] == '{': current_obj_type = parts[1] brace_stack.append('{') current_obj_content.append(line) continue # 处理行内的括号 if '{' in stripped_line: brace_stack.extend(['{' for _ in stripped_line.split('{') if _]) if '}' in stripped_line: for _ in stripped_line.split('}'): if brace_stack: brace_stack.pop() # 栈空=当前对象结束 if not brace_stack: current_obj_content.append(line) # 执行你的类型分支逻辑 if current_obj_type == 'node': print(f"Processing node content:\n{''.join(current_obj_content)}") # 这里替换成你的to do 1逻辑 elif current_obj_type == 'triplex_meter': print(f"Processing triplex_meter content:\n{''.join(current_obj_content)}") # 这里替换成你的to do 2逻辑 elif current_obj_type == 'triplex_line': print(f"Processing triplex_line content:\n{''.join(current_obj_content)}") # 这里替换成对应的处理逻辑 # 重置状态,准备下一个对象 current_obj_type = None current_obj_content = [] else: current_obj_content.append(line) continue # 如果当前处于对象内部,收集内容行 if current_obj_type is not None: current_obj_content.append(line)
方法二:正则表达式(适合无嵌套的简单场景)
如果你的文件里没有嵌套的大括号(看你提到有嵌套,这个方法就不太适用,但如果大部分对象都是平级的可以试试),可以用正则一次性匹配整个对象块,代码会更简洁:
import re def parse_with_regex(file_path): with open(file_path, 'r') as f: content = f.read() # 匹配所有object块,注意:嵌套大括号会导致匹配失败 pattern = re.compile(r'object (\w+) \{([\s\S]+?)\}', re.MULTILINE) matches = pattern.findall(content) for obj_type, obj_content in matches: # 清理内容:去掉注释和空行 cleaned_content = '\n'.join([ line.strip() for line in obj_content.split('\n') if line.strip() and not line.strip().startswith('#') ]) if obj_type == 'node': print(f"Processing node:\n{cleaned_content}") # 替换成你的to do 1逻辑 elif obj_type == 'triplex_meter': print(f"Processing triplex_meter:\n{cleaned_content}") # 替换成你的to do 2逻辑 elif obj_type == 'triplex_line': print(f"Processing triplex_line:\n{cleaned_content}") # 替换成对应的处理逻辑
⚠️ 注意:这个正则用了非贪婪匹配[\s\S]+?,但如果对象内部有嵌套的{},比如某个属性值里带大括号,正则会提前匹配到第一个}就结束,导致内容截断,所以嵌套场景还是优先用栈的方法。
额外优化建议
- 如果文件特别大,一定要用逐行读取的栈方法,别一次性读入整个文件,避免内存占用过高
- 可以把对象内容进一步解析成字典,比如把
name R2-12-47-3_node_453;转成{'name': 'R2-12-47-3_node_453'},后续处理属性会更方便:
def parse_obj_content(content_lines): obj_dict = {} for line in content_lines: stripped = line.strip() if not stripped or stripped.startswith('#') or '{' in stripped or '}' in stripped: continue # 分割属性名和值,支持值带空格的情况(用maxsplit=1) parts = stripped.rstrip(';').split(maxsplit=1) if len(parts) == 2: key, value = parts obj_dict[key] = value return obj_dict
然后在栈方法里,当对象结束时调用这个函数,就能拿到结构化的字典,处理逻辑会清晰很多。
内容的提问来源于stack exchange,提问作者Xu Siyuan

