如何用Python解析层级混合文本文件并生成表格?
层级文本解析与字段填充解决方案
核心思路
通过维护上下文状态跟踪当前所在的组(Group)和章节(Section),逐行解析文本的缩进层级,自动为每条数据填充对应的组/章节信息,最终整理成目标表格结构。
具体实现步骤
- 预处理每行文本:逐行读取文件,计算每行开头的空格缩进数,同时去除行尾换行符和首尾多余空格,分离缩进层级与内容。
- 维护上下文变量:定义
current_group和current_section两个变量,分别存储当前所处的组和章节信息,初始值为空。 - 按层级分支处理:
- 缩进0空格:文件首行,无需解析可直接跳过(若有特殊需求可单独处理)。
- 缩进2空格:识别为Group行,更新
current_group为当前行内容,同时重置current_section和临时数据存储字典(新组下无默认章节)。 - 缩进4空格:识别为Section行,更新
current_section为当前行内容,重置临时数据存储字典。 - 缩进6空格:识别为数据标签(如Number、Date、Time),记录当前标签类型,为后续填充值做准备。
- 缩进8空格:识别为数据值,将值存入临时字典对应标签的键下;当临时字典集齐
Number、Date、Time三个字段时,补充当前current_group和current_section信息,将字典加入结果列表,再重置临时字典以处理下一条数据。
Python代码示例
def parse_hierarchical_file(file_path): current_group = "" current_section = "" current_field = "" result = [] temp_data = {} # 逐行读取,避免加载大型文件占用过多内存 with open(file_path, 'r', encoding='utf-8') as f: for line in f: line = line.rstrip('\n') # 计算缩进空格数 indent = len(line) - len(line.lstrip(' ')) content = line.strip() if indent == 0: continue elif indent == 2: current_group = content current_section = "" temp_data = {} elif indent == 4: current_section = content temp_data = {} elif indent == 6: # 假设标签格式为"Number:",提取标签名 current_field = content.split(':')[0].strip() elif indent == 8: temp_data[current_field] = content # 检查是否收集完所有必填字段 if all(key in temp_data for key in ['Number', 'Date', 'Time']): temp_data['Group'] = current_group temp_data['Section'] = current_section result.append(temp_data.copy()) temp_data = {} return result # 调用示例,转换为表格(需安装pandas) if __name__ == "__main__": parsed_data = parse_hierarchical_file("target_file.txt") import pandas as pd df = pd.DataFrame(parsed_data) print(df)
注意事项
- 若数据标签与值的对应逻辑有变化(如标签不含冒号),需调整标签提取的代码逻辑。
- 针对超大型文件,可进一步优化为边解析边写入输出文件(如CSV),避免内存积压。
- 若组/章节名称包含固定格式的编号(如
Group-001),可使用正则表达式(如r'Group-(\d+)')提取精准编号,替代直接存储完整名称。
内容的提问来源于stack exchange,提问作者CardPlayer
相关产品推荐
相关产品推荐

