如何将带分段的TXT文件拆分为以子标题为键的Pandas DataFrame字典?
按子标题拆分TXT文件并生成DataFrame字典
针对你的需求,这里提供直接可运行的解决方案,核心思路是通过状态遍历识别子标题和表格内容,再将表格转换为Pandas DataFrame存入字典:
import pandas as pd # 初始化存储结果的字典和临时变量 df_dict = {} current_subtitle = None is_collecting_table = False table_content = [] with open('filename.txt', encoding='utf8') as f: for line in f: cleaned_line = line.strip() # 跳过空行和主标题行 if not cleaned_line or cleaned_line == 'Title': continue # 识别子标题行,处理上一个表格(如果存在) if cleaned_line.startswith('Sub Heading:'): # 若已有未处理的表格数据,先转成DataFrame存入字典 if current_subtitle and table_content: # 第一行作为列名,后续行作为数据 df = pd.DataFrame(table_content[1:], columns=table_content[0]) # 将字符串数值转为数值类型 df = df.apply(pd.to_numeric) df_dict[current_subtitle] = df table_content = [] # 提取子标题(如A、B、C) current_subtitle = cleaned_line.split(':')[-1].strip() continue # 切换表格收集状态(遇到```时) if cleaned_line == '```': is_collecting_table = not is_collecting_table continue # 收集表格行数据(仅在收集状态且非空行时) if is_collecting_table and cleaned_line: # 按任意数量空格分割列,自动处理表格的空格分隔 table_content.append(cleaned_line.split()) # 处理最后一个子标题对应的表格 if current_subtitle and table_content: df = pd.DataFrame(table_content[1:], columns=table_content[0]) df = df.apply(pd.to_numeric) df_dict[current_subtitle] = df # 验证结果(可选) for subtitle, df in df_dict.items(): print(f"=== {subtitle} 对应的DataFrame ===") print(df) print("\n")
关键细节说明
- 状态控制:用
is_collecting_table标记是否处于表格内容收集阶段,遇到代码块标记\```时切换状态。 - 子标题切换处理:每次识别到新的子标题时,先将上一个子标题对应的表格数据转换为DataFrame,避免数据遗漏。
- 表格解析:用
split()处理空格分隔的表格行,自动适配多个空格的分隔场景;第一行作为列名,后续行作为数据行。 - 数据类型转换:通过
apply(pd.to_numeric)将字符串格式的数值转为数值类型,确保后续数据分析的可用性。
内容的提问来源于stack exchange,提问作者CB_datarookie
相关产品推荐
相关产品推荐

