如何逐行读取文件匹配字符串并将结果拆分至多个文件
Got it, let's fix this and get your "chicken" lines split across multiple files! First, a quick note on your original code: opening the output file in a+ mode every time you write a line is pretty inefficient, especially for large input files. We'll fix that while adding the split functionality.
Below are two practical approaches to split your matching lines into multiple files—pick the one that fits your needs:
1. Split by Fixed Number of Lines
If you want to group, say, every 10 "chicken" lines into a separate file (like output_1.txt, output_2.txt, etc.), use this method:
def split_lines_by_count(filename, new_filename_prefix, lines_per_file=10): line_count = 0 file_index = 1 current_output = None try: with open(filename, 'r') as rl: for line in rl: if "chicken" in line: # Open a new file if we haven't started one, or hit the line limit if current_output is None or line_count >= lines_per_file: # Close the previous file if it exists if current_output: current_output.close() # Generate the new filename (e.g., output_1.txt) current_filename = f"{new_filename_prefix}_{file_index}.txt" current_output = open(current_filename, 'w') file_index += 1 line_count = 0 # Reset counter for the new file current_output.write(line) line_count += 1 finally: # Make sure we close the last open file even if something goes wrong if current_output: current_output.close() # Example usage: Split every 5 matching lines into a new file split_lines_by_count("input.txt", "output", lines_per_file=5)
How this works:
- We track how many lines we've written to the current output file with
line_count. - When we hit the
lines_per_filelimit, we close the current file, increment the file index, and open a new one. - The
finallyblock ensures we don't leave any file handles hanging if an error occurs.
2. Split by File Size
If you want to split files once they reach a certain size (like 1MB), use this approach instead:
import os def split_lines_by_size(filename, new_filename_prefix, max_size_bytes=1024*1024): # 1MB default file_index = 1 current_output = None current_size = 0 try: with open(filename, 'r') as rl: for line in rl: if "chicken" in line: # Calculate the actual byte size of the line (string length != file size for special chars) line_size = len(line.encode('utf-8')) # Check if adding this line would exceed the max size, or if we need a new file if current_output is None or (current_size + line_size) > max_size_bytes: # Close previous file if current_output: current_output.close() # Create new file current_filename = f"{new_filename_prefix}_{file_index}.txt" current_output = open(current_filename, 'w') file_index += 1 current_size = 0 current_output.write(line) current_size += line_size finally: if current_output: current_output.close() # Example usage: Split when files reach 500KB split_lines_by_size("input.txt", "output", max_size_bytes=500*1024)
How this works:
- We calculate the byte size of each line to accurately track file growth (since plain string length doesn't account for multi-byte characters).
- When adding the next line would push the current file over
max_size_bytes, we switch to a new file.
Bonus: Split by Custom Logic
If you need to split based on specific patterns in your input (like starting a new file every time you hit a ## comment ## line), you can add a custom check to either approach. For example:
# Add this inside the loop where you process lines if "## comment ##" in line: # Close current file and start a new one if current_output: current_output.close() current_output = open(f"{new_filename_prefix}_{file_index}.txt", 'w') file_index += 1 line_count = 0 # or current_size = 0, depending on your base approach
Just tweak the condition to match whatever custom rule you need!
内容的提问来源于stack exchange,提问作者Izzat Z.

