Python处理12GB大文本时list引发内存溢出,该如何修改程序?
Python处理12GB大文件内存溢出问题修复方案
核心原因
问题本质是将1.5亿行的拆分结果全部存入内存级的Python列表data_seq_arr和data_tag_arr,列表的元数据+存储的字符串内容总大小超过了系统可分配的内存上限,必然会触发OOM崩溃。
解决方案
按实际使用场景选择对应方案即可:
方案1:流式逐行处理(最推荐,内存占用最低)
如果不需要一次性访问全量的seq和tag数据,直接把后续的处理逻辑嵌入到逐行读取的循环内,完全不需要定义全局列表存储所有结果。
示例代码:
with open("150_million.txt", "r", encoding="utf-8") as f1: for line in f1: line_arr = line.split("@@@") if len(line_arr) != 2: continue seq, tag = line_arr if " " not in seq or " " not in tag: continue seq_arr = seq.split(" ") tag_arr = tag.split(" ") # 这里直接写当前行的处理逻辑,比如写入结果文件、计算指标、送入模型推理等 process_single_line(seq_arr, tag_arr)
方案2:分批次处理
如果后续逻辑需要批量处理数据(比如批量喂入模型训练),设置批次大小,攒够指定行数后处理一次,再清空列表释放内存:
示例代码:
BATCH_SIZE = 10000 # 可根据内存大小调整 batch_seq = [] batch_tag = [] with open("150_million.txt", "r", encoding="utf-8") as f1: for line in f1: line_arr = line.split("@@@") if len(line_arr) != 2: continue seq, tag = line_arr if " " not in seq or " " not in tag: continue seq_arr = seq.split(" ") tag_arr = tag.split(" ") batch_seq.append(seq_arr) batch_tag.append(tag_arr) # 攒够批次就处理,然后清空 if len(batch_seq) >= BATCH_SIZE: process_batch(batch_seq, batch_tag) batch_seq.clear() batch_tag.clear() # 处理最后不足一批的剩余数据 if batch_seq: process_batch(batch_seq, batch_tag)
方案3:全量数据持久化存储
如果业务必须要随机访问全量的seq、tag数据,不要用内存列表存储,将结果写入本地持久化存储介质:
- 小规模随机访问可以用
sqlite数据库存储每行的seq和tag,需要的时候按行号查询 - 数值类数据可以用
numpy.memmap做内存映射,直接操作磁盘文件,读写逻辑和内存数组一致,不会占满内存 - 结构化数据可以用
pandas的to_parquet分块写入,后续按块读取
额外优化点
- 读文件统一用
with上下文管理器,避免文件句柄泄漏 - 如果seq、tag都是数值类型,转成
numpy数组存储比原生Python列表内存占用低60%以上 - 如果需要对外提供可迭代的全量数据接口,用生成器替代列表,惰性加载不会占用额外内存,示例:
def iter_seq_tag(file_path: str): with open(file_path, "r", encoding="utf-8") as f: for line in f: line_arr = line.split("@@@") if len(line_arr) != 2: continue seq, tag = line_arr if " " not in seq or " " not in tag: continue yield seq.split(" "), tag.split(" ") # 使用时直接遍历即可,不会占用内存存储全量数据 for seq_arr, tag_arr in iter_seq_tag("150_million.txt"): do_something(seq_arr, tag_arr)
内容的提问来源于stack exchange,提问作者user9799714
相关产品推荐
相关产品推荐

