Pandas读取缩进分隔文本文件构建多列DataFrame实现方案
缩进文本转结构化DataFrame实现方案
字段逻辑分类
先根据字段的生效范围和出现规律分类,避免跨行填充逻辑混乱:
- 作用域级公共字段:Id、set、Main_id、Secondary_id、Start_Date、End_Date,这类字段在所属Id/set范围内全局有效,所有该范围内生成的数据行都复用对应值,解析到新值时直接覆盖旧值即可
- 行级分组字段:Quantity、Type、Value、Category、Capacity,这类字段按组连续出现,每凑齐一组(以
Capacity作为每组结束标识,样例中每组均以Category+Capacity收尾)就对应一行独立数据
解析流程设计
不需要靠统计缩进数量判断层级,直接通过字段规律和作用域切换逻辑即可完成解析:
- 文本预处理:逐行读取文件,跳过空行、开头的总计数行,对每行去除首尾空白,按第一个冒号拆分为键、值两部分,统一去除键值前后的多余空格
- 分层缓存设计:
- 顶层变量存当前解析到的Id、set编号,遇到新Id/set时先触发上一个set的收尾逻辑,再更新值
- set级缓存存当前set下的Main_id、Secondary_id、Start_Date、End_Date值
- 待处理行列表暂存当前set下已经解析完成的分组行(因为Start_Date、End_Date固定出现在set末尾,早于这两个字段解析到的分组行先存在列表里,等set收尾时统一补全字段)
- 行临时缓存存当前正在拼接的分组字段,每遇到Capacity字段,就把临时缓存的内容存入待处理行列表,清空临时缓存准备解析下一行
- set收尾逻辑:遇到新Id、新set、文件读取结束时,遍历待处理行列表,给每一行补上当前作用域的所有公共字段,存入最终结果集,清空当前set的缓存
可运行代码
import pandas as pd # 定义最终输出的列顺序 q_columns = ["Id", "set","Main_id", "Secondary_id","Quantity", "Type", "Value", "Category", "Capacity", "Start_Date", "End_Date"] # 初始化各层缓存 current_id = None current_set = None set_common = {} pending_rows = [] current_row = {} result = [] def flush_pending(): """set解析结束时,批量补全公共字段写入最终结果""" nonlocal pending_rows, set_common, current_id, current_set, result for row in pending_rows: full_row = { "Id": current_id, "set": current_set, **set_common, **row } result.append({col: full_row.get(col) for col in q_columns}) pending_rows = [] set_common = {} with open('test.txt', 'r', encoding='utf-8') as f: for line in f: line = line.strip() # 跳过无效行 if not line or line.startswith('Total ids entered'): continue # 拆分键值对,仅按第一个冒号拆分,兼容值中包含冒号的场景 key, val = line.split(':', 1) key = key.strip() val = val.strip() if key == "Id": # 切换Id前先处理完上一个set的待存行 if current_set is not None: flush_pending() current_id = val current_set = None elif key == "set": # 切换set前先处理完上一个set的待存行 if current_set is not None: flush_pending() current_set = val elif key in {"Main_id", "Secondary_id", "Start_Date", "End_Date"}: set_common[key] = val elif key in {"Quantity", "Type", "Value", "Category", "Capacity"}: current_row[key] = val # 遇到Capacity标记当前分组行拼接完成,加入待处理列表 if key == "Capacity": pending_rows.append(current_row.copy()) current_row = {} # 文件读取结束,处理最后一个set的待存行 if current_set is not None: flush_pending() # 生成目标DataFrame df_q = pd.DataFrame(result, columns=q_columns)
适配说明
- 代码完全匹配提供的样例结构:Id=6050下set=256的2组Category/Capacity会拆分为2行,Id=123下set=789的2组不同Quantity/Type/Category组合会拆分为2行,所有上层公共字段会自动填充
- 不需要依赖缩进计数,即使后续文本缩进格式有小幅调整,只要字段名和出现顺序规律不变就能正常解析
- 自动处理字段缺失场景,若某个字段未解析到会填充为空值,不会报错中断
内容的提问来源于stack exchange,提问作者user7675621
相关产品推荐
相关产品推荐

