You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python文件解析与层级嵌套数据结构构建需求

Python文件解析与层级嵌套数据结构构建需求

嘿,看起来你需要处理一个带有多层嵌套结构的配置文件,要把它转换成Python里好用的层级数据结构对吧?我刚好做过类似的需求,给你梳理下思路和具体实现方案。

首先先明确下文件的结构规则,咱们先把它理清楚:

  • 整个文件包含一个或多个plan块,每个plan以plan [名称]开头,endplan或endplan //[名称]结束
  • 每个plan下面可以有多个feature块,feature同样以feature [名称]开头,endfeature //[名称]结束,而且feature还能嵌套子feature,形成多层结构
  • 每个feature下面可以有多个measure块,measure以measure [名称] :开头,endmeasure //[名称]结束,里面会包含src = "..."这类属性字段

数据结构设计思路

既然是多层嵌套,用字典+列表的组合来构建数据结构是最直观的,每个层级的对象都用字典存储(包含名称和子元素),子元素用列表来维护顺序:

  • 最外层是一个plans列表,每个元素是一个plan字典
  • Plan字典:包含name(plan名称)和features列表(存储该plan下的所有feature)
  • Feature字典:包含name(feature名称)、measures列表(存储该feature下的measure)、sub_features列表(存储嵌套的子feature)
  • Measure字典:包含name(measure名称)和src(对应的src属性值,后续有其他属性可以直接扩展字段)

具体解析代码实现

核心思路是用**栈(Stack)**来处理嵌套结构——因为嵌套块是后进先出的,进入一个块就把它压入栈,结束时弹出,这样就能始终跟踪当前正在处理的层级。我们用正则表达式来匹配每行的格式,确保精准解析:

import re
from pprint import pprint

def parse_plan_file(file_path):
    plans = []
    stack = []
    
    # 定义各类型行的正则匹配模式
    plan_pattern = re.compile(r'^plan (\w+)$')
    feature_pattern = re.compile(r'^feature (\w+)$')
    measure_start_pattern = re.compile(r'^measure (\w+) :$')
    src_pattern = re.compile(r'^src = "(.*)"$')
    endmeasure_pattern = re.compile(r'^endmeasure //(\w+)$')
    endfeature_pattern = re.compile(r'^endfeature // (\w+)$')
    endplan_pattern = re.compile(r'^endplan( //(\w+))?$')

    with open(file_path, 'r', encoding='utf-8') as f:
        for line_num, line in enumerate(f, 1):
            line = line.strip()
            # 跳过空行和纯注释行
            if not line or line.startswith('//'):
                continue
            # 移除行内的注释部分(比如//后面的内容)
            line = line.split('//')[0].strip()
            if not line:
                continue

            # 处理plan开头
            plan_match = plan_pattern.match(line)
            if plan_match:
                plan_name = plan_match.group(1)
                new_plan = {
                    'name': plan_name,
                    'features': []
                }
                plans.append(new_plan)
                stack.append(new_plan)
                continue

            # 处理feature开头
            feature_match = feature_pattern.match(line)
            if feature_match:
                feature_name = feature_match.group(1)
                new_feature = {
                    'name': feature_name,
                    'measures': [],
                    'sub_features': []
                }
                # 判断当前父级是plan还是feature,添加到对应的列表
                current_parent = stack[-1]
                if 'features' in current_parent:
                    current_parent['features'].append(new_feature)
                else:
                    current_parent['sub_features'].append(new_feature)
                stack.append(new_feature)
                continue

            # 处理measure开头
            measure_match = measure_start_pattern.match(line)
            if measure_match:
                measure_name = measure_match.group(1)
                new_measure = {
                    'name': measure_name,
                    'src': None
                }
                # 父级必须是feature
                if stack and 'measures' in stack[-1]:
                    stack[-1]['measures'].append(new_measure)
                    stack.append(new_measure)
                continue

            # 处理src属性行
            src_match = src_pattern.match(line)
            if src_match:
                src_value = src_match.group(1)
                # 当前层级必须是measure
                if stack and 'src' in stack[-1]:
                    stack[-1]['src'] = src_value
                continue

            # 处理measure结束
            endmeasure_match = endmeasure_pattern.match(line)
            if endmeasure_match:
                measure_name = endmeasure_match.group(1)
                # 校验结束名称和当前measure是否匹配(可选,用于格式校验)
                if stack and stack[-1]['name'] == measure_name:
                    stack.pop()
                continue

            # 处理feature结束
            endfeature_match = endfeature_pattern.match(line)
            if endfeature_match:
                feature_name = endfeature_match.group(1)
                if stack and stack[-1]['name'] == feature_name:
                    stack.pop()
                continue

            # 处理plan结束
            endplan_match = endplan_pattern.match(line)
            if endplan_match:
                plan_name = endplan_match.group(2)
                if stack:
                    if plan_name and stack[-1]['name'] != plan_name:
                        print(f"⚠️ 警告:第{line_num}行,endplan名称{plan_name}与当前plan{stack[-1]['name']}不匹配")
                    stack.pop()
                continue

            # 未匹配到任何模式的行,输出警告
            print(f"⚠️ 警告:第{line_num}行,无法识别的内容:{line}")

    return plans

# 测试示例
if __name__ == "__main__":
    parsed_plans = parse_plan_file("your_plan_file.txt")
    print("解析结果:")
    pprint(parsed_plans)

代码关键点说明

  1. 栈的使用:完美解决嵌套层级的跟踪问题,确保每个子块都能正确归属到父块下
  2. 正则匹配:精准匹配每行格式,避免误解析;同时处理了行内注释的情况
  3. 格式校验:加入了简单的名称匹配校验,能及时发现文件中的格式错误
  4. 可扩展性:如果后续需要添加measure的其他属性(比如type、name),只需要新增对应的正则模式和字典字段即可

扩展建议

  • 如果需要更面向对象的管理,可以把Plan、Feature、Measure封装成类,用实例来存储数据,后续操作会更灵活
  • 如果处理超大文件,可以考虑用生成器逐块解析,避免一次性加载整个文件到内存
  • 可以添加更严格的错误处理(比如抛出异常),而不是只打印警告,让程序对格式错误更敏感

备注:内容来源于stack exchange,提问作者Alok

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.21 11:58:07