You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何逐行读取文件匹配字符串并将结果拆分至多个文件

Split Matching Lines into Multiple Files

Got it, let's fix this and get your "chicken" lines split across multiple files! First, a quick note on your original code: opening the output file in a+ mode every time you write a line is pretty inefficient, especially for large input files. We'll fix that while adding the split functionality.

Below are two practical approaches to split your matching lines into multiple files—pick the one that fits your needs:

1. Split by Fixed Number of Lines

If you want to group, say, every 10 "chicken" lines into a separate file (like output_1.txt, output_2.txt, etc.), use this method:

def split_lines_by_count(filename, new_filename_prefix, lines_per_file=10):
    line_count = 0
    file_index = 1
    current_output = None

    try:
        with open(filename, 'r') as rl:
            for line in rl:
                if "chicken" in line:
                    # Open a new file if we haven't started one, or hit the line limit
                    if current_output is None or line_count >= lines_per_file:
                        # Close the previous file if it exists
                        if current_output:
                            current_output.close()
                        # Generate the new filename (e.g., output_1.txt)
                        current_filename = f"{new_filename_prefix}_{file_index}.txt"
                        current_output = open(current_filename, 'w')
                        file_index += 1
                        line_count = 0  # Reset counter for the new file

                    current_output.write(line)
                    line_count += 1
    finally:
        # Make sure we close the last open file even if something goes wrong
        if current_output:
            current_output.close()

# Example usage: Split every 5 matching lines into a new file
split_lines_by_count("input.txt", "output", lines_per_file=5)

How this works:

  • We track how many lines we've written to the current output file with line_count.
  • When we hit the lines_per_file limit, we close the current file, increment the file index, and open a new one.
  • The finally block ensures we don't leave any file handles hanging if an error occurs.

2. Split by File Size

If you want to split files once they reach a certain size (like 1MB), use this approach instead:

import os

def split_lines_by_size(filename, new_filename_prefix, max_size_bytes=1024*1024):  # 1MB default
    file_index = 1
    current_output = None
    current_size = 0

    try:
        with open(filename, 'r') as rl:
            for line in rl:
                if "chicken" in line:
                    # Calculate the actual byte size of the line (string length != file size for special chars)
                    line_size = len(line.encode('utf-8'))
                    # Check if adding this line would exceed the max size, or if we need a new file
                    if current_output is None or (current_size + line_size) > max_size_bytes:
                        # Close previous file
                        if current_output:
                            current_output.close()
                        # Create new file
                        current_filename = f"{new_filename_prefix}_{file_index}.txt"
                        current_output = open(current_filename, 'w')
                        file_index += 1
                        current_size = 0

                    current_output.write(line)
                    current_size += line_size
    finally:
        if current_output:
            current_output.close()

# Example usage: Split when files reach 500KB
split_lines_by_size("input.txt", "output", max_size_bytes=500*1024)

How this works:

  • We calculate the byte size of each line to accurately track file growth (since plain string length doesn't account for multi-byte characters).
  • When adding the next line would push the current file over max_size_bytes, we switch to a new file.

Bonus: Split by Custom Logic

If you need to split based on specific patterns in your input (like starting a new file every time you hit a ## comment ## line), you can add a custom check to either approach. For example:

# Add this inside the loop where you process lines
if "## comment ##" in line:
    # Close current file and start a new one
    if current_output:
        current_output.close()
    current_output = open(f"{new_filename_prefix}_{file_index}.txt", 'w')
    file_index += 1
    line_count = 0  # or current_size = 0, depending on your base approach

Just tweak the condition to match whatever custom rule you need!

内容的提问来源于stack exchange,提问作者Izzat Z.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:50:56