You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python处理12GB大文本时list引发内存溢出,该如何修改程序?

Python处理12GB大文件内存溢出问题修复方案

核心原因

问题本质是将1.5亿行的拆分结果全部存入内存级的Python列表data_seq_arr和data_tag_arr,列表的元数据+存储的字符串内容总大小超过了系统可分配的内存上限,必然会触发OOM崩溃。

解决方案

按实际使用场景选择对应方案即可:

方案1:流式逐行处理(最推荐,内存占用最低)

如果不需要一次性访问全量的seq和tag数据,直接把后续的处理逻辑嵌入到逐行读取的循环内,完全不需要定义全局列表存储所有结果。
示例代码:

with open("150_million.txt", "r", encoding="utf-8") as f1:
    for line in f1:
        line_arr = line.split("@@@")
        if len(line_arr) != 2:
            continue
        seq, tag = line_arr
        if " " not in seq or " " not in tag:
            continue
        seq_arr = seq.split(" ")
        tag_arr = tag.split(" ")
        # 这里直接写当前行的处理逻辑,比如写入结果文件、计算指标、送入模型推理等
        process_single_line(seq_arr, tag_arr)

方案2:分批次处理

如果后续逻辑需要批量处理数据(比如批量喂入模型训练),设置批次大小,攒够指定行数后处理一次,再清空列表释放内存:
示例代码:

BATCH_SIZE = 10000 # 可根据内存大小调整
batch_seq = []
batch_tag = []
with open("150_million.txt", "r", encoding="utf-8") as f1:
    for line in f1:
        line_arr = line.split("@@@")
        if len(line_arr) != 2:
            continue
        seq, tag = line_arr
        if " " not in seq or " " not in tag:
            continue
        seq_arr = seq.split(" ")
        tag_arr = tag.split(" ")
        batch_seq.append(seq_arr)
        batch_tag.append(tag_arr)
        # 攒够批次就处理,然后清空
        if len(batch_seq) >= BATCH_SIZE:
            process_batch(batch_seq, batch_tag)
            batch_seq.clear()
            batch_tag.clear()
    # 处理最后不足一批的剩余数据
    if batch_seq:
        process_batch(batch_seq, batch_tag)

方案3:全量数据持久化存储

如果业务必须要随机访问全量的seq、tag数据,不要用内存列表存储,将结果写入本地持久化存储介质:

  • 小规模随机访问可以用sqlite数据库存储每行的seq和tag,需要的时候按行号查询
  • 数值类数据可以用numpy.memmap做内存映射,直接操作磁盘文件,读写逻辑和内存数组一致,不会占满内存
  • 结构化数据可以用pandas的to_parquet分块写入,后续按块读取

额外优化点

  • 读文件统一用with上下文管理器,避免文件句柄泄漏
  • 如果seq、tag都是数值类型,转成numpy数组存储比原生Python列表内存占用低60%以上
  • 如果需要对外提供可迭代的全量数据接口,用生成器替代列表,惰性加载不会占用额外内存,示例:
def iter_seq_tag(file_path: str):
    with open(file_path, "r", encoding="utf-8") as f:
        for line in f:
            line_arr = line.split("@@@")
            if len(line_arr) != 2:
                continue
            seq, tag = line_arr
            if " " not in seq or " " not in tag:
                continue
            yield seq.split(" "), tag.split(" ")

# 使用时直接遍历即可,不会占用内存存储全量数据
for seq_arr, tag_arr in iter_seq_tag("150_million.txt"):
    do_something(seq_arr, tag_arr)

内容的提问来源于stack exchange,提问作者user9799714

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 23:09:02