You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆分TXT多块数据为独立dataframe并合并为带分类字段的总表

实现方案

实现思路

直接按行遍历原始TXT文件完成两个需求:

  • 检测到Test_Type开头的行就判定为新数据块开始,完成上一个块的拆分存储后初始化新块的处理流程
  • 每个块先解析键: 值格式的头部公共属性,空行后解析数据区的列名和实际数据,给所有数据行关联上当前块的头部属性,同时完成单独存储和合并准备

完整代码

import pandas as pd
import os

# 配置路径参数
raw_file_path = "替换为你的原始TXT文件路径"
split_output_dir = "./split_test_data"  # 拆分后单个csv的存储目录
os.makedirs(split_output_dir, exist_ok=True)

# 初始化处理变量
current_header = {}
current_data = []
in_data_zone = False
data_cols = []
merged_df_list = []

# 逐行读取原始文件
with open(raw_file_path, 'r', encoding='utf-8') as f:
    for line in f:
        line_content = line.strip()
        # 触发新块开始逻辑
        if line_content.startswith("Test_Type"):
            # 先处理完上一个块的残留数据
            if current_data and current_header:
                # 转DataFrame并关联头部属性
                block_df = pd.DataFrame(current_data, columns=data_cols)
                for attr_key, attr_val in current_header.items():
                    block_df[attr_key] = attr_val
                # 保存单独csv
                csv_file_name = f"{current_header['Test_Type']}_Cond{current_header['Condition_Number']}_Trial{current_header['Trial_Number']}.csv"
                block_df.to_csv(os.path.join(split_output_dir, csv_file_name), index=False)
                # 加入合并列表
                merged_df_list.append(block_df)
            # 重置变量处理新块
            current_header = {}
            current_data = []
            in_data_zone = False
            data_cols = []
            # 解析Test_Type属性
            key, val = line_content.split(":", 1)
            current_header[key.strip()] = val.strip()
            continue
        # 解析头部属性行
        if ":" in line_content and not in_data_zone:
            key, val = line_content.split(":", 1)
            current_header[key.strip()] = val.strip()
            continue
        # 空行判定头部结束,进入数据区
        if not line_content and not in_data_zone:
            in_data_zone = True
            continue
        # 处理数据区内容
        if in_data_zone and line_content:
            line_split = line_content.split()
            # 跳过单位行
            if line_split[0] == "UNITS":
                continue
            # 读取数据列名
            if line_split[0] == "DP":
                data_cols = line_split
                continue
            # 存储数据行(转浮点型)
            current_data.append([float(x) for x in line_split])

# 处理文件末尾最后一个未处理的数据块
if current_data and current_header:
    block_df = pd.DataFrame(current_data, columns=data_cols)
    for attr_key, attr_val in current_header.items():
        block_df[attr_key] = attr_val
    csv_file_name = f"{current_header['Test_Type']}_Cond{current_header['Condition_Number']}_Trial{current_header['Trial_Number']}.csv"
    block_df.to_csv(os.path.join(split_output_dir, csv_file_name), index=False)
    merged_df_list.append(block_df)

# 合并所有块得到最终总DataFrame
total_merged_df = pd.concat(merged_df_list, ignore_index=True)
# 可按需保存总文件
total_merged_df.to_csv("all_test_data_merged.csv", index=False)

注意事项

  • 若原始文件编码报错,可把open函数里的encoding='utf-8'替换为encoding='gbk'或者encoding='latin-1'
  • 若需要保留单位信息,可删除跳过UNITS行的逻辑,把单位信息也存入头部属性字典
  • 若数据行存在非数值内容,可调整最后一行的数值转换逻辑适配你的实际数据格式

内容的提问来源于stack exchange,提问作者Snowbowl18

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 20:45:09