You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将定长TXT合并为CSV时遇MemoryError的高效解决方法问询

内存友好的定长TXT文件合并方案(适配16GB内存限制)

1. 先优化单文件读取的内存占用

用read_fwf时,pandas默认的类型推断会浪费不少内存,手动指定数据类型能大幅降低单文件的内存占用:

  • 给每列指定最小合适的数据类型:比如用int8/int16替代默认的int64,float32替代float64;如果某列字符串重复值多,直接设为category类型
  • 只读取需要的列:用usecols参数过滤掉无关列,减少不必要的内存消耗
  • 关闭缺失值检测:如果确认文件里没有缺失值,加na_filter=False参数,能省点内存和读取时间

示例代码:

import pandas as pd

# 自定义列名(根据你的定长字段数量调整)
col_names = [f"col_{i}" for i in range(10)]  # 假设你有10个字段
# 自定义各列数据类型,根据实际数据调整
dtype_map = {
    "col_0": "int32",
    "col_1": "float32",
    "col_2": "category",  # 该列重复值多的话用这个
    "col_3": "int16"
}

# 优化后的单文件读取
df = pd.read_fwf(
    "your_file.txt",
    header=None,
    names=col_names,
    dtype=dtype_map,
    usecols=[0,1,2,3],  # 只读取前4列,按需调整
    na_filter=False
)

2. 边读边写,别一次性加载所有文件

别把所有文件的DataFrame都存到列表里再合并,这样内存会直接爆掉。改成读一个文件就追加到目标文件,内存里始终只留单个文件(或分块)的数据:

方案A:追加到CSV文件

CSV支持直接追加,适合边读边写:

import os
import pandas as pd

file_dir = "你的文件存放目录"
all_txt_files = [f for f in os.listdir(file_dir) if f.endswith(".txt")]
col_names = [f"col_{i}" for i in range(10)]
dtype_map = {"col_0": "int32", "col_1": "float32"}  # 沿用之前的类型定义

# 先处理第一个文件,写入完整内容(包含表头)
first_file_path = os.path.join(file_dir, all_txt_files[0])
df_first = pd.read_fwf(first_file_path, header=None, names=col_names, dtype=dtype_map)
df_first.to_csv("combined_data.csv", index=False)

# 循环处理剩下的文件,只追加数据(不写表头)
for file_name in all_txt_files[1:]:
    file_path = os.path.join(file_dir, file_name)
    df = pd.read_fwf(file_path, header=None, names=col_names, dtype=dtype_map)
    df.to_csv("combined_data.csv", mode="a", header=False, index=False)
    # 手动释放内存,针对大文件更有用
    del df

方案B:转成Pickle文件(更快的读取速度)

Pickle不支持直接追加,所以可以先合并成CSV,最后再转成Pickle:

# 先执行方案A生成combined_data.csv,再执行以下代码转Pickle
df_combined = pd.read_csv("combined_data.csv", dtype=dtype_map)
df_combined.to_pickle("combined_data.pkl")

3. 超大文件(1GB级)分块读取

单个1GB的文件直接读可能也会占太多内存,用chunksize分块读取,逐块追加:

import pandas as pd

large_file_path = "large_1gb_file.txt"
col_names = [f"col_{i}" for i in range(10)]
dtype_map = {"col_0": "int32", "col_1": "float32"}

# 生成分块迭代器,每次读10万行(可根据内存调整)
chunk_iter = pd.read_fwf(
    large_file_path,
    header=None,
    names=col_names,
    dtype=dtype_map,
    chunksize=100000
)

# 处理第一个块,写入文件
first_chunk = next(chunk_iter)
first_chunk.to_csv("large_file_combined.csv", index=False)

# 处理剩余块,逐块追加
for chunk in chunk_iter:
    chunk.to_csv("large_file_combined.csv", mode="a", header=False, index=False)
    del chunk

4. 额外的内存优化小技巧

  • 加low_memory=False:读取时加上这个参数,避免pandas为了省内存对列类型做分段推断,防止后续合并时出现类型不一致的问题
  • 手动清理内存:循环里处理完单个文件/分块后,用del df和gc.collect()强制释放内存(尤其处理大文件时):
import gc

for file_name in all_txt_files[1:]:
    file_path = os.path.join(file_dir, file_name)
    df = pd.read_fwf(file_path, header=None, names=col_names, dtype=dtype_map)
    df.to_csv("combined_data.csv", mode="a", header=False, index=False)
    del df
    gc.collect()

内容的提问来源于stack exchange,提问作者PSt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 17:55:13