如何用Python/Pandas读取含多表的制表符分隔文本文件生成多DataFrame?
读取特殊格式文本文件生成多DataFrame(Python/Pandas实现)
实现方案
直接逐行解析文件,根据标记识别表格结构,提取表头和数据后转换为DataFrame:
import pandas as pd def parse_formatted_file(file_path): table_dict = {} current_table_id = None current_headers = [] current_rows = [] with open(file_path, 'r', encoding='utf-8') as f: for raw_line in f: line = raw_line.strip() if not line: continue # 触发文件结束,终止解析 if line.startswith('%E'): break # 新表格起始 if line.startswith('%T'): # 先保存上一个已收集完的表格 if current_table_id and current_headers and current_rows: table_dict[current_table_id] = pd.DataFrame(current_rows, columns=current_headers) # 提取表格名称,无名称则自动生成 parts = line.split('\t') current_table_id = parts[1] if len(parts) > 1 else f"table_{len(table_dict)+1}" # 重置当前表格的表头和数据行 current_headers = [] current_rows = [] # 提取列标题 elif line.startswith('%F'): current_headers = line.split('\t')[1:] # 提取数据记录 elif line.startswith('%R'): row_parts = line.split('\t')[1:] # 校验列数匹配,避免脏数据混入 if len(row_parts) == len(current_headers): current_rows.append(row_parts) # 处理最后一个未保存的表格 if current_table_id and current_headers and current_rows: table_dict[current_table_id] = pd.DataFrame(current_rows, columns=current_headers) return table_dict # 调用示例 target_file = "your_data_file.txt" result_tables = parse_formatted_file(target_file) # 查看结果 for name, df in result_tables.items(): print(f"\n--- 表格 {name} ---") print(df.head())
关键细节说明
- 自动跳过无关内容:开头的无关文本不会匹配任何标记(%T/%F/%R/%E),会被直接跳过。
- 表格命名逻辑:如果
%T行后有制表符分隔的自定义名称,就用该名称;否则自动生成顺序命名(如table_1)。 - 数据校验:仅保留列数与表头一致的记录行,避免因脏数据导致DataFrame结构错误。
- 终止机制:遇到
%E标记立即停止解析,无需处理后续内容。
内容的提问来源于stack exchange,提问作者Matthew
相关产品推荐
相关产品推荐

