You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas按指定分隔字符串拆分文本为多个DataFrame?

问题

我有如下格式的文本文件(以空格分隔数据):

LINE 1 to SKIP
LINE 2 to SKIP
2.13999987 0.139999986 -0.398405492 1
2.61999989 6.0000062E-2 0.450082362 1
2.74000001 5.99999428E-2 1.04403841 1
2.84000015 4.00000811E-2 6.17375337E-2 1
IGN IGN IGN IGN 
21.4200001 0.420000076 1.53572667 1
22.3199997 0.479999542 -0.595370948 1
23.3199997 0.520000458 0.136062101 1
24.3600006 0.519999504 -0.520044923 1
25.3999996 0.520000458 2.45230961 1
26.4399986 0.519999504 -2.08248448 1
27.4799995 0.520000458 -0.263438225 1
IGN IGN IGN IGN 
58.6800003 0.520000458 -0.789233088 1
59.7200012 0.520000458 -1.02961564 1
60.7600021 0.51999855 -0.889572859 1
61.7999992 0.520000458 -1.03346229 1
62.8400002 0.520000458 4.94940579E-2 1

目前我能通过以下代码读取单个数据块:

df_first = pd.read_table('file.txt', names=names, delimiter=' ', skiprows=3, nrows=4)

其中names是各列的名称。我需要把每个连续的数据行块分配给指定名称的DataFrame(可通过名称数组指定),遇到IGN IGN IGN IGN分隔行就切换到下一个DataFrame,直到文件结束。请问最优实现方法是什么?

最优实现方法

可以通过逐行读取文件拆分数据块,再批量转换为DataFrame,既节省内存又能精准控制数据分配,具体实现如下:

代码实现

import pandas as pd

# 配置参数
file_path = 'file.txt'
names = ['col1', 'col2', 'col3', 'col4']  # 替换为你的实际列名
df_names = ['df1', 'df2', 'df3']  # 指定每个数据块对应的DataFrame名称
separator = 'IGN IGN IGN IGN'

# 初始化数据块收集器
data_blocks = []
current_block = []

# 逐行读取并拆分数据
with open(file_path, 'r') as f:
    # 跳过开头两行无效内容
    next(f)
    next(f)
    for line in f:
        line = line.strip()
        if not line:
            continue
        # 遇到分隔符时,保存当前数据块并重置收集器
        if line == separator:
            if current_block:
                data_blocks.append(current_block)
                current_block = []
        else:
            # 处理连续空格分隔的情况,转换数据类型
            row_data = line.split()
            row = [float(val) for val in row_data]
            current_block.append(row)
    # 处理文件末尾最后一个未被分隔符结尾的数据块
    if current_block:
        data_blocks.append(current_block)

# 将数据块转为DataFrame,存入字典按名称调用
df_dict = {}
for df_name, block in zip(df_names, data_blocks):
    df_dict[df_name] = pd.DataFrame(block, columns=names)

# 示例:查看第一个DataFrame
print(df_dict['df1'])

关键说明

  • 逐行读取避免一次性加载大文件,适合处理超大型数据集
  • 自动处理连续空格分隔的格式问题,兼容文件中的数据样式
  • 通过字典存储DataFrame,可直接通过指定名称快速访问对应数据块
  • 若DataFrame名称数量与数据块数量不匹配,代码会自动截断多余块(可根据需求调整逻辑)

内容的提问来源于stack exchange,提问作者Py-ser

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 22:02:12