You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas重构含重复分段结构DataFrame的实现方法

通用CSV分段数据重构方案

核心逻辑是通过固定的SECTION起始标记切分所有分段,不依赖单段固定行数,适配任意数量分段的解析需求。


实现步骤

  • 读取原始CSV时不预设全局表头,按原始行顺序读入所有内容,避免把分段内部的属性行、迭代表头误识别为全局表头。
  • 遍历所有行定位分段边界:找到第一列值为SECTION的所有行索引,额外补全文件开头第一个无SECTION标记的分段起始位置(对应示例里最开头的LOT1分段),再把文件末尾位置补为最后一个分段的结束边界。
  • 逐块处理切分好的分段:
    1. 先清理块内的全空无效行
    2. 从块首开始逐行提取固定属性:LOT、DESCRIPTION、MEAN、MIN、MAX,直到遇到值为Iteration #的行停止
    3. 以Iteration #所在行为子表头,定位需要提取的Value、A、B列位置,逐行提取后续所有迭代轮次的对应字段值,按顺序拼接到固定属性后面
  • 把所有分段处理得到的等长序列按列拼接,生成最终DataFrame,每一列对应一个原始分段。

可直接运行的代码

import pandas as pd
import numpy as np

# 读取原始文件,按多空格/制表符分隔,所有内容按字符串读入,空值填充为空串
raw_df = pd.read_csv(
    "your_file_path.csv",
    header=None,
    sep=r"\s+",  # 如果是逗号分隔就改成sep=","
    dtype=str
).fillna("")

# 定位所有分段的起始、结束索引
section_starts = list(raw_df[raw_df[0].str.strip() == "SECTION"].index)
# 补全文件开头没有SECTION标记的首段
if 0 not in section_starts:
    section_starts = [0] + section_starts
# 追加文件末尾作为最后一段的结束边界
section_starts.append(len(raw_df))

processed_sections = []
for seg_idx in range(len(section_starts) - 1):
    s, e = section_starts[seg_idx], section_starts[seg_idx+1]
    seg_block = raw_df.iloc[s:e].reset_index(drop=True)
    # 过滤全空行
    seg_block = seg_block[~seg_block.apply(lambda x: x.str.strip().eq("").all(), axis=1)].reset_index(drop=True)

    seg_data = []
    cursor = 0
    # 提取固定属性
    while cursor < len(seg_block):
        first_col_val = seg_block.iloc[cursor, 0].strip()
        if first_col_val == "Iteration #":
            break
        # 取该行最后一个非空值为属性值
        attr_val = seg_block.iloc[cursor].replace("", np.nan).dropna().iloc[-1].strip()
        seg_data.append(attr_val)
        cursor += 1

    # 定位迭代数据的列索引
    iter_header = [i.strip() for i in seg_block.iloc[cursor].replace("", np.nan).dropna().tolist()]
    # 注意第一列是Iteration编号,列索引需要偏移1位
    val_col_pos = iter_header.index("Value") + 1
    a_col_pos = iter_header.index("A") + 1
    b_col_pos = iter_header.index("B") + 1

    # 提取所有迭代轮次数据
    cursor += 1
    while cursor < len(seg_block):
        row = seg_block.iloc[cursor]
        seg_data.extend([
            row.iloc[val_col_pos].strip(),
            row.iloc[a_col_pos].strip(),
            row.iloc[b_col_pos].strip()
        ])
        cursor += 1
    processed_sections.append(seg_data)

# 生成最终结果表
final_df = pd.DataFrame(processed_sections).T
# 配置行索引名称,可按需修改
row_labels = ["LOT", "DESCRIPTION", "MEAN", "MIN", "MAX"]
iter_count = (len(final_df) - 5) // 3
for round_num in range(iter_count):
    row_labels.extend([
        f"Iter{round_num+1}_Value",
        f"Iter{round_num+1}_A",
        f"Iter{round_num+1}_B"
    ])
final_df.index = row_labels

适配说明

该方案不限制单分段的行数、不限制分段总数量,只要所有分段结构一致、以SECTION为起始标识、固定属性在前且迭代数据表头为Iteration #即可正常解析。如果需要提取其他字段,直接修改属性提取、迭代列定位的对应逻辑即可。

内容的提问来源于stack exchange,提问作者Jan Rostenkowski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 07:48:06