You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python解析层级混合文本文件并生成表格?

层级文本解析与字段填充解决方案

核心思路

通过维护上下文状态跟踪当前所在的组(Group)和章节(Section),逐行解析文本的缩进层级,自动为每条数据填充对应的组/章节信息,最终整理成目标表格结构。


具体实现步骤

  • 预处理每行文本:逐行读取文件,计算每行开头的空格缩进数,同时去除行尾换行符和首尾多余空格,分离缩进层级与内容。
  • 维护上下文变量:定义current_group和current_section两个变量,分别存储当前所处的组和章节信息,初始值为空。
  • 按层级分支处理:
    • 缩进0空格:文件首行,无需解析可直接跳过(若有特殊需求可单独处理)。
    • 缩进2空格:识别为Group行,更新current_group为当前行内容,同时重置current_section和临时数据存储字典(新组下无默认章节)。
    • 缩进4空格:识别为Section行,更新current_section为当前行内容,重置临时数据存储字典。
    • 缩进6空格:识别为数据标签(如Number、Date、Time),记录当前标签类型,为后续填充值做准备。
    • 缩进8空格:识别为数据值,将值存入临时字典对应标签的键下;当临时字典集齐Number、Date、Time三个字段时,补充当前current_group和current_section信息,将字典加入结果列表,再重置临时字典以处理下一条数据。

Python代码示例

def parse_hierarchical_file(file_path):
    current_group = ""
    current_section = ""
    current_field = ""
    result = []
    temp_data = {}

    # 逐行读取,避免加载大型文件占用过多内存
    with open(file_path, 'r', encoding='utf-8') as f:
        for line in f:
            line = line.rstrip('\n')
            # 计算缩进空格数
            indent = len(line) - len(line.lstrip(' '))
            content = line.strip()

            if indent == 0:
                continue
            elif indent == 2:
                current_group = content
                current_section = ""
                temp_data = {}
            elif indent == 4:
                current_section = content
                temp_data = {}
            elif indent == 6:
                # 假设标签格式为"Number:",提取标签名
                current_field = content.split(':')[0].strip()
            elif indent == 8:
                temp_data[current_field] = content
                # 检查是否收集完所有必填字段
                if all(key in temp_data for key in ['Number', 'Date', 'Time']):
                    temp_data['Group'] = current_group
                    temp_data['Section'] = current_section
                    result.append(temp_data.copy())
                    temp_data = {}
    return result

# 调用示例,转换为表格(需安装pandas)
if __name__ == "__main__":
    parsed_data = parse_hierarchical_file("target_file.txt")
    import pandas as pd
    df = pd.DataFrame(parsed_data)
    print(df)

注意事项

  • 若数据标签与值的对应逻辑有变化(如标签不含冒号),需调整标签提取的代码逻辑。
  • 针对超大型文件,可进一步优化为边解析边写入输出文件(如CSV),避免内存积压。
  • 若组/章节名称包含固定格式的编号(如Group-001),可使用正则表达式(如r'Group-(\d+)')提取精准编号,替代直接存储完整名称。

内容的提问来源于stack exchange,提问作者CardPlayer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 15:28:22