Python逐行读取大文件:满足条件时跳过对应行数的实现
Solution for Processing Large Files Line-by-Line
Since your file is a whopping 7.2GB, loading the entire thing into a list is going to eat up way too much memory—so a line-by-line approach is exactly what you need. Here's a memory-efficient implementation that only keeps track of the current line and how many lines we need to skip next:
The Code
def process_large_file(input_file_path, output_file_path=None): skip_count = 0 # Open the input file (and output file if specified) with open(input_file_path, 'r') as infile: # Handle the header line (remove this block if your file has no header) header_line = next(infile) if output_file_path: with open(output_file_path, 'w') as outfile: outfile.write(header_line) _process_lines(infile, outfile, skip_count) else: # If no output file, just print the results (adjust as needed) print(header_line.strip()) _process_lines(infile, None, skip_count) def _process_lines(infile, outfile, skip_count): for line in infile: # If we need to skip lines, just decrement the counter and move on if skip_count > 0: skip_count -= 1 continue # Parse the line to get Col2 value line_parts = line.strip().split() if len(line_parts) < 2: # Skip malformed lines that don't have two columns continue try: col2_value = int(line_parts[1]) except ValueError: # Handle cases where Col2 isn't an integer if outfile: outfile.write(line) else: print(line.strip()) continue # Keep the current line (write to file or print) if outfile: outfile.write(line) else: print(line.strip()) # If Col2 is greater than 1, set the number of lines to skip next if col2_value > 1: skip_count = col2_value
How It Works
- Memory Efficiency: We only load one line into memory at a time, plus a tiny counter (
skip_count). This means memory usage stays constant no matter how big your file is. - Skip Logic: When we hit a line where Col2 > 1, we set
skip_countto that value. Every subsequent line just decrements this counter until it hits 0, at which point we start processing lines again. - Error Handling: The code includes checks for malformed lines (missing columns, non-integer Col2 values) to avoid crashes mid-processing.
- Flexibility: You can either print the results to the console or write them directly to a new file by passing an
output_file_path.
Example Usage
To process your file and save the results to a new file:
process_large_file("your_large_file.txt", "processed_output.txt")
To just print the processed lines to the console:
process_large_file("your_large_file.txt")
This approach will handle your 7.2GB file without breaking a sweat, no more memory overload issues!
内容的提问来源于stack exchange,提问作者Homap
相关产品推荐
相关产品推荐

