Python文件解析与层级嵌套数据结构构建需求
Python文件解析与层级嵌套数据结构构建需求
嘿,看起来你需要处理一个带有多层嵌套结构的配置文件,要把它转换成Python里好用的层级数据结构对吧?我刚好做过类似的需求,给你梳理下思路和具体实现方案。
首先先明确下文件的结构规则,咱们先把它理清楚:
- 整个文件包含一个或多个plan块,每个plan以
plan [名称]开头,endplan或endplan //[名称]结束 - 每个plan下面可以有多个feature块,feature同样以
feature [名称]开头,endfeature //[名称]结束,而且feature还能嵌套子feature,形成多层结构 - 每个feature下面可以有多个measure块,measure以
measure [名称] :开头,endmeasure //[名称]结束,里面会包含src = "..."这类属性字段
数据结构设计思路
既然是多层嵌套,用字典+列表的组合来构建数据结构是最直观的,每个层级的对象都用字典存储(包含名称和子元素),子元素用列表来维护顺序:
- 最外层是一个
plans列表,每个元素是一个plan字典 - Plan字典:包含
name(plan名称)和features列表(存储该plan下的所有feature) - Feature字典:包含
name(feature名称)、measures列表(存储该feature下的measure)、sub_features列表(存储嵌套的子feature) - Measure字典:包含
name(measure名称)和src(对应的src属性值,后续有其他属性可以直接扩展字段)
具体解析代码实现
核心思路是用**栈(Stack)**来处理嵌套结构——因为嵌套块是后进先出的,进入一个块就把它压入栈,结束时弹出,这样就能始终跟踪当前正在处理的层级。我们用正则表达式来匹配每行的格式,确保精准解析:
import re from pprint import pprint def parse_plan_file(file_path): plans = [] stack = [] # 定义各类型行的正则匹配模式 plan_pattern = re.compile(r'^plan (\w+)$') feature_pattern = re.compile(r'^feature (\w+)$') measure_start_pattern = re.compile(r'^measure (\w+) :$') src_pattern = re.compile(r'^src = "(.*)"$') endmeasure_pattern = re.compile(r'^endmeasure //(\w+)$') endfeature_pattern = re.compile(r'^endfeature // (\w+)$') endplan_pattern = re.compile(r'^endplan( //(\w+))?$') with open(file_path, 'r', encoding='utf-8') as f: for line_num, line in enumerate(f, 1): line = line.strip() # 跳过空行和纯注释行 if not line or line.startswith('//'): continue # 移除行内的注释部分(比如//后面的内容) line = line.split('//')[0].strip() if not line: continue # 处理plan开头 plan_match = plan_pattern.match(line) if plan_match: plan_name = plan_match.group(1) new_plan = { 'name': plan_name, 'features': [] } plans.append(new_plan) stack.append(new_plan) continue # 处理feature开头 feature_match = feature_pattern.match(line) if feature_match: feature_name = feature_match.group(1) new_feature = { 'name': feature_name, 'measures': [], 'sub_features': [] } # 判断当前父级是plan还是feature,添加到对应的列表 current_parent = stack[-1] if 'features' in current_parent: current_parent['features'].append(new_feature) else: current_parent['sub_features'].append(new_feature) stack.append(new_feature) continue # 处理measure开头 measure_match = measure_start_pattern.match(line) if measure_match: measure_name = measure_match.group(1) new_measure = { 'name': measure_name, 'src': None } # 父级必须是feature if stack and 'measures' in stack[-1]: stack[-1]['measures'].append(new_measure) stack.append(new_measure) continue # 处理src属性行 src_match = src_pattern.match(line) if src_match: src_value = src_match.group(1) # 当前层级必须是measure if stack and 'src' in stack[-1]: stack[-1]['src'] = src_value continue # 处理measure结束 endmeasure_match = endmeasure_pattern.match(line) if endmeasure_match: measure_name = endmeasure_match.group(1) # 校验结束名称和当前measure是否匹配(可选,用于格式校验) if stack and stack[-1]['name'] == measure_name: stack.pop() continue # 处理feature结束 endfeature_match = endfeature_pattern.match(line) if endfeature_match: feature_name = endfeature_match.group(1) if stack and stack[-1]['name'] == feature_name: stack.pop() continue # 处理plan结束 endplan_match = endplan_pattern.match(line) if endplan_match: plan_name = endplan_match.group(2) if stack: if plan_name and stack[-1]['name'] != plan_name: print(f"⚠️ 警告:第{line_num}行,endplan名称{plan_name}与当前plan{stack[-1]['name']}不匹配") stack.pop() continue # 未匹配到任何模式的行,输出警告 print(f"⚠️ 警告:第{line_num}行,无法识别的内容:{line}") return plans # 测试示例 if __name__ == "__main__": parsed_plans = parse_plan_file("your_plan_file.txt") print("解析结果:") pprint(parsed_plans)
代码关键点说明
- 栈的使用:完美解决嵌套层级的跟踪问题,确保每个子块都能正确归属到父块下
- 正则匹配:精准匹配每行格式,避免误解析;同时处理了行内注释的情况
- 格式校验:加入了简单的名称匹配校验,能及时发现文件中的格式错误
- 可扩展性:如果后续需要添加measure的其他属性(比如type、name),只需要新增对应的正则模式和字典字段即可
扩展建议
- 如果需要更面向对象的管理,可以把Plan、Feature、Measure封装成类,用实例来存储数据,后续操作会更灵活
- 如果处理超大文件,可以考虑用生成器逐块解析,避免一次性加载整个文件到内存
- 可以添加更严格的错误处理(比如抛出异常),而不是只打印警告,让程序对格式错误更敏感
备注:内容来源于stack exchange,提问作者Alok
相关产品推荐
相关产品推荐

