You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效读取多份文本文件至带ID列的pandas DataFrame

问题描述

我有多份如下格式的文本文件:

对应ID A的文件内容:

#-----------------------------------------
#    foo  bar  baz  
#-----------------------------------------
    0.0120932       1.10166       1.08745 
    0.0127890       1.10105       1.08773 
    0.0142051       1.09941       1.08760 
    0.0162801       1.09662       1.08548 
    0.0197376       1.09170       1.08015 

对应ID B的文件内容:

#-----------------------------------------
#    foo  bar  baz  
#-----------------------------------------
    0.888085      0.768590      0.747961
    0.893782      0.781607      0.760417
    0.899830      0.797021      0.771219
    0.899266      0.799260      0.765859
    0.891489      0.781255      0.728892

需要将所有文件读取到列名为 ['ID', 'foo', 'bar', 'baz'] 的pandas DataFrame中,文件名与ID的对应关系通过字典映射(例如file_id_map = {"file_A.txt": "A", "file_B.txt": "B"})。

之前尝试过逐行读取转字典再生成DataFrame,但这种方式不适合单文件多数据行的场景;拼接多个小DataFrame的效率又太低,求更高效的解决方案。

期望输出示例:

ID       foo      bar       baz
0   A  0.012093  1.10166   1.08745
1   A  0.012789  1.10105   1.08773
2   A  0.014205  1.09941   1.08760
3   A  0.016280  1.09662   1.08548
4   A  0.019738  1.09170   1.08015
5   B  0.888085  0.768590  0.747961
6   B  0.893782  0.781607  0.760417
7   B  0.899830  0.797021  0.771219
8   B  0.899266  0.799260  0.765859
9   B  0.891489  0.781255  0.728892
高效解决方案

方法1:批量读取+一次性合并(简洁高效)

利用pandas.read_csv的参数跳过注释行直接读取数据,给每个文件的DataFrame添加ID列后,一次性合并所有结果,比逐行处理或多次拼接效率高很多。

import pandas as pd

# 文件名到ID的映射字典
file_id_map = {
    "file_A.txt": "A",
    "file_B.txt": "B"
    # 添加其他文件映射
}

# 批量处理生成带ID的DataFrame列表
dfs = []
for file_path, id_val in file_id_map.items():
    # skiprows跳过前3行(空行+分隔线+表头注释行),sep匹配任意空白符
    df = pd.read_csv(
        file_path,
        skiprows=3,
        sep=r"\s+",
        names=["foo", "bar", "baz"],
        dtype=float
    )
    df["ID"] = id_val
    # 调整列顺序,将ID置于首位
    df = df[["ID", "foo", "bar", "baz"]]
    dfs.append(df)

# 一次性合并并重置索引
final_df = pd.concat(dfs, ignore_index=True)
print(final_df)

方法2:生成器优化内存(超大量文件适用)

如果文件数量极多、数据量庞大,用生成器替代列表存储中间DataFrame,减少内存占用:

import pandas as pd

file_id_map = {"file_A.txt": "A", "file_B.txt": "B"}

def read_file_generator(file_map):
    for file_path, id_val in file_map.items():
        df = pd.read_csv(
            file_path,
            skiprows=3,
            sep=r"\s+",
            names=["foo", "bar", "baz"]
        )
        df["ID"] = id_val
        yield df[["ID", "foo", "bar", "baz"]]

# 从生成器合并数据
final_df = pd.concat(read_file_generator(file_id_map), ignore_index=True)

方法3:直接读取数据行创建DataFrame(极致高效)

追求最高效率的话,可直接读取所有文件的有效数据行,整理成列表后一次性创建DataFrame,避免多次创建小DataFrame的开销:

import pandas as pd

file_id_map = {"file_A.txt": "A", "file_B.txt": "B"}

data = []
for file_path, id_val in file_id_map.items():
    with open(file_path, "r") as f:
        # 跳过前3行无效内容
        for _ in range(3):
            next(f)
        # 读取每一行有效数据
        for line in f:
            line = line.strip()
            if not line:
                continue
            foo, bar, baz = map(float, line.split())
            data.append([id_val, foo, bar, baz])

# 直接从列表生成DataFrame
final_df = pd.DataFrame(data, columns=["ID", "foo", "bar", "baz"])
print(final_df)

方法对比

  • 方法1:兼顾代码简洁性和效率,pandas内部优化过读取逻辑,适合绝大多数场景。
  • 方法2:适合超大量文件/数据,生成器按需返回DataFrame,降低内存压力。
  • 方法3:效率最高,避免pandas读取的额外开销,但代码稍繁琐,适合对性能要求极高的场景。

内容的提问来源于stack exchange,提问作者DeltaIV

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 05:05:21