You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手求助:如何低内存打开1GB级gzip文件

Fixing Memory Issues When Reading Large Gzip Files in Python

Hey there! I see you're hitting a memory error while trying to process that 1GB gzip file, and I can help you sort this out. Let's break down what's going wrong and how to fix it properly.

First, Why Your Current Code Is Failing

Your code has two key issues:

  1. When you loop with for line in peak, line is already a single line of bytes from the gzip file. Calling line.read() on a bytes object doesn't make sense—that's exactly why you're getting that error from the gzip.py module.
  2. Even if that didn't throw an error, trying to load the entire 1GB file into memory at once would crash due to memory limits.

The Low-Memory Solution: Process the File Line-by-Line

The trick to handling huge files without eating up all your memory is to process them one line at a time. The gzip.open() object is built to be iterable, meaning it gives you one line at a time automatically—no need to load everything into memory upfront.

Here's the corrected, memory-friendly code:

import gzip

# Open the gzip file in read-binary mode
with gzip.open("test.gz", "rb") as peak:
    # Iterate over each line (as bytes) directly from the file
    for line_bytes in peak:
        # Decode the bytes to a UTF-8 string
        line_content = line_bytes.decode("utf-8")
        # Print or process the line here
        print(line_content)

Extra Tips for Optimizing Further

If you only need to extract specific parts of the file (not every line), add a condition to skip lines you don't care about—this saves even more memory and processing time:

import gzip

# Example: Only keep lines containing a specific keyword
target_keyword = "critical_info"

with gzip.open("test.gz", "rb") as peak:
    for line_bytes in peak:
        line_content = line_bytes.decode("utf-8")
        if target_keyword in line_content:
            # Process or save the relevant line
            print("Found relevant content:", line_content)
            # To save to a file instead of printing:
            # with open("extracted_content.txt", "a", encoding="utf-8") as out_file:
            #     out_file.write(line_content)

Why This Works

By iterating directly over the gzip.open() object, you only keep one line of the file in memory at a time. This uses a tiny amount of memory, regardless of how big the original gzip file is (even if it's 10GB+!).

内容的提问来源于stack exchange,提问作者Gyan Prakash Mishra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:46:30