You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python逐行读取大文件:满足条件时跳过对应行数的实现

Solution for Processing Large Files Line-by-Line

Since your file is a whopping 7.2GB, loading the entire thing into a list is going to eat up way too much memory—so a line-by-line approach is exactly what you need. Here's a memory-efficient implementation that only keeps track of the current line and how many lines we need to skip next:

The Code

def process_large_file(input_file_path, output_file_path=None):
    skip_count = 0
    
    # Open the input file (and output file if specified)
    with open(input_file_path, 'r') as infile:
        # Handle the header line (remove this block if your file has no header)
        header_line = next(infile)
        if output_file_path:
            with open(output_file_path, 'w') as outfile:
                outfile.write(header_line)
                _process_lines(infile, outfile, skip_count)
        else:
            # If no output file, just print the results (adjust as needed)
            print(header_line.strip())
            _process_lines(infile, None, skip_count)

def _process_lines(infile, outfile, skip_count):
    for line in infile:
        # If we need to skip lines, just decrement the counter and move on
        if skip_count > 0:
            skip_count -= 1
            continue
        
        # Parse the line to get Col2 value
        line_parts = line.strip().split()
        if len(line_parts) < 2:
            # Skip malformed lines that don't have two columns
            continue
        
        try:
            col2_value = int(line_parts[1])
        except ValueError:
            # Handle cases where Col2 isn't an integer
            if outfile:
                outfile.write(line)
            else:
                print(line.strip())
            continue
        
        # Keep the current line (write to file or print)
        if outfile:
            outfile.write(line)
        else:
            print(line.strip())
        
        # If Col2 is greater than 1, set the number of lines to skip next
        if col2_value > 1:
            skip_count = col2_value

How It Works

  • Memory Efficiency: We only load one line into memory at a time, plus a tiny counter (skip_count). This means memory usage stays constant no matter how big your file is.
  • Skip Logic: When we hit a line where Col2 > 1, we set skip_count to that value. Every subsequent line just decrements this counter until it hits 0, at which point we start processing lines again.
  • Error Handling: The code includes checks for malformed lines (missing columns, non-integer Col2 values) to avoid crashes mid-processing.
  • Flexibility: You can either print the results to the console or write them directly to a new file by passing an output_file_path.

Example Usage

To process your file and save the results to a new file:

process_large_file("your_large_file.txt", "processed_output.txt")

To just print the processed lines to the console:

process_large_file("your_large_file.txt")

This approach will handle your 7.2GB file without breaking a sweat, no more memory overload issues!

内容的提问来源于stack exchange,提问作者Homap

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:45:42