You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python/Pandas读取含多表的制表符分隔文本文件生成多DataFrame?

读取特殊格式文本文件生成多DataFrame(Python/Pandas实现)

实现方案

直接逐行解析文件,根据标记识别表格结构,提取表头和数据后转换为DataFrame:

import pandas as pd

def parse_formatted_file(file_path):
    table_dict = {}
    current_table_id = None
    current_headers = []
    current_rows = []

    with open(file_path, 'r', encoding='utf-8') as f:
        for raw_line in f:
            line = raw_line.strip()
            if not line:
                continue
            
            # 触发文件结束,终止解析
            if line.startswith('%E'):
                break
            
            # 新表格起始
            if line.startswith('%T'):
                # 先保存上一个已收集完的表格
                if current_table_id and current_headers and current_rows:
                    table_dict[current_table_id] = pd.DataFrame(current_rows, columns=current_headers)
                # 提取表格名称,无名称则自动生成
                parts = line.split('\t')
                current_table_id = parts[1] if len(parts) > 1 else f"table_{len(table_dict)+1}"
                # 重置当前表格的表头和数据行
                current_headers = []
                current_rows = []
            
            # 提取列标题
            elif line.startswith('%F'):
                current_headers = line.split('\t')[1:]
            
            # 提取数据记录
            elif line.startswith('%R'):
                row_parts = line.split('\t')[1:]
                # 校验列数匹配,避免脏数据混入
                if len(row_parts) == len(current_headers):
                    current_rows.append(row_parts)
        
        # 处理最后一个未保存的表格
        if current_table_id and current_headers and current_rows:
            table_dict[current_table_id] = pd.DataFrame(current_rows, columns=current_headers)
    
    return table_dict

# 调用示例
target_file = "your_data_file.txt"
result_tables = parse_formatted_file(target_file)

# 查看结果
for name, df in result_tables.items():
    print(f"\n--- 表格 {name} ---")
    print(df.head())

关键细节说明

  • 自动跳过无关内容:开头的无关文本不会匹配任何标记(%T/%F/%R/%E),会被直接跳过。
  • 表格命名逻辑:如果%T行后有制表符分隔的自定义名称,就用该名称;否则自动生成顺序命名(如table_1)。
  • 数据校验:仅保留列数与表头一致的记录行,避免因脏数据导致DataFrame结构错误。
  • 终止机制:遇到%E标记立即停止解析,无需处理后续内容。

内容的提问来源于stack exchange,提问作者Matthew

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 22:10:03