You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas读取缩进分隔文本文件构建多列DataFrame实现方案

缩进文本转结构化DataFrame实现方案

字段逻辑分类

先根据字段的生效范围和出现规律分类,避免跨行填充逻辑混乱:

  • 作用域级公共字段:Id、set、Main_id、Secondary_id、Start_Date、End_Date,这类字段在所属Id/set范围内全局有效,所有该范围内生成的数据行都复用对应值,解析到新值时直接覆盖旧值即可
  • 行级分组字段:Quantity、Type、Value、Category、Capacity,这类字段按组连续出现,每凑齐一组(以Capacity作为每组结束标识,样例中每组均以Category+Capacity收尾)就对应一行独立数据

解析流程设计

不需要靠统计缩进数量判断层级,直接通过字段规律和作用域切换逻辑即可完成解析:

  1. 文本预处理:逐行读取文件,跳过空行、开头的总计数行,对每行去除首尾空白,按第一个冒号拆分为键、值两部分,统一去除键值前后的多余空格
  2. 分层缓存设计:
    • 顶层变量存当前解析到的Id、set编号,遇到新Id/set时先触发上一个set的收尾逻辑,再更新值
    • set级缓存存当前set下的Main_id、Secondary_id、Start_Date、End_Date值
    • 待处理行列表暂存当前set下已经解析完成的分组行(因为Start_Date、End_Date固定出现在set末尾,早于这两个字段解析到的分组行先存在列表里,等set收尾时统一补全字段)
    • 行临时缓存存当前正在拼接的分组字段,每遇到Capacity字段,就把临时缓存的内容存入待处理行列表,清空临时缓存准备解析下一行
  3. set收尾逻辑:遇到新Id、新set、文件读取结束时,遍历待处理行列表,给每一行补上当前作用域的所有公共字段,存入最终结果集,清空当前set的缓存

可运行代码

import pandas as pd

# 定义最终输出的列顺序
q_columns = ["Id", "set","Main_id", "Secondary_id","Quantity", "Type",
       "Value", "Category", "Capacity", "Start_Date", "End_Date"]

# 初始化各层缓存
current_id = None
current_set = None
set_common = {}
pending_rows = []
current_row = {}
result = []

def flush_pending():
    """set解析结束时,批量补全公共字段写入最终结果"""
    nonlocal pending_rows, set_common, current_id, current_set, result
    for row in pending_rows:
        full_row = {
            "Id": current_id,
            "set": current_set,
            **set_common,
            **row
        }
        result.append({col: full_row.get(col) for col in q_columns})
    pending_rows = []
    set_common = {}

with open('test.txt', 'r', encoding='utf-8') as f:
    for line in f:
        line = line.strip()
        # 跳过无效行
        if not line or line.startswith('Total ids entered'):
            continue
        # 拆分键值对,仅按第一个冒号拆分,兼容值中包含冒号的场景
        key, val = line.split(':', 1)
        key = key.strip()
        val = val.strip()

        if key == "Id":
            # 切换Id前先处理完上一个set的待存行
            if current_set is not None:
                flush_pending()
            current_id = val
            current_set = None
        elif key == "set":
            # 切换set前先处理完上一个set的待存行
            if current_set is not None:
                flush_pending()
            current_set = val
        elif key in {"Main_id", "Secondary_id", "Start_Date", "End_Date"}:
            set_common[key] = val
        elif key in {"Quantity", "Type", "Value", "Category", "Capacity"}:
            current_row[key] = val
            # 遇到Capacity标记当前分组行拼接完成,加入待处理列表
            if key == "Capacity":
                pending_rows.append(current_row.copy())
                current_row = {}
    # 文件读取结束,处理最后一个set的待存行
    if current_set is not None:
        flush_pending()

# 生成目标DataFrame
df_q = pd.DataFrame(result, columns=q_columns)

适配说明

  • 代码完全匹配提供的样例结构:Id=6050下set=256的2组Category/Capacity会拆分为2行,Id=123下set=789的2组不同Quantity/Type/Category组合会拆分为2行,所有上层公共字段会自动填充
  • 不需要依赖缩进计数,即使后续文本缩进格式有小幅调整,只要字段名和出现顺序规律不变就能正常解析
  • 自动处理字段缺失场景,若某个字段未解析到会填充为空值,不会报错中断

内容的提问来源于stack exchange,提问作者user7675621

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 05:18:19