You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取大型压缩/非压缩文件内存过高进程被杀死如何优化

问题核心原因
  • 内存溢出的直接诱因是你调用了vcf.readlines():这个方法会一次性把整个文件的所有内容读入内存,拼装成列表返回,超大文件(尤其是基因测序场景下的大体积VCF文件)解压后全部加载必然会占满内存,导致进程被系统OOM机制杀死。
  • 额外隐藏问题:你调用gzip.open()时没有指定文本模式,默认返回的是二进制读模式,读取到的line是bytes类型,直接和字符串'#'做匹配会抛出类型错误,你现在还没遇到这个报错是因为进程先被内存问题杀掉了。
修复方案

首先把迭代逻辑改成直接遍历文件对象本身,文件对象是原生可迭代对象,会逐行读取内容,每次只加载一行到内存,完全避免一次性加载全量文件的问题:

def process_vcf(location):
    logging.info('Processing vcf')
    logging.debug(location)
    with read_compressed_or_not(location) as vcf:
        # 去掉readlines(),直接迭代文件对象
        for line in vcf:
            if line.startswith('#'):
                logging.debug(line)

然后调整上下文管理器里的gzip.open调用,指定文本模式,和普通文本文件的读取行为统一:

@contextmanager
def read_compressed_or_not(location):
    if location.endswith('.gz'):
        try: 
            # 指定rt模式即文本读模式,自动解码bytes为字符串
            file = gzip.open(location, 'rt')
            yield file
        finally:
            file.close()
    else:
        try: 
            file = open(location, 'r')
            yield file
        finally:
            file.close()
可选优化

如果你的Python版本 >= 3.7,上下文管理器可以简化写法,不需要自己手写try-finally逻辑,因为gzip.open和普通open本身就实现了上下文管理协议,可直接嵌套with语句自动处理资源释放:

from contextlib import contextmanager

@contextmanager
def read_compressed_or_not(location):
    if location.endswith('.gz'):
        with gzip.open(location, 'rt') as f:
            yield f
    else:
        with open(location, 'r') as f:
            yield f

内容的提问来源于stack exchange,提问作者Moopsish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 22:45:05