You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

超2亿行大文本批量替换:用首块10-100行替换所有区块对应行

Solution for Replacing Specific Ranges Across Chunks in a Large Text File

Got it, let's break this down—handling a 200M-line file sounds daunting, but by leaning into chunked processing and lightweight tools, we can get this done efficiently without blowing up your memory. Here are two solid approaches depending on your comfort with command-line tools vs. scripting:

Approach 1: Command-Line Toolchain (Fast, No Scripting)

This uses split, head, tail, and cat—all standard Unix/Linux tools that process files in streams (so no need to load the entire 200M-line file into memory).

Step 1: Extract the template lines from the first chunk

First, grab lines 10-100 from the first 50k-line chunk and save them to a temporary template file:

head -n 100 big_file.txt | tail -n 91 > template_lines.txt

(91 because 100-10+1 = 91 lines total)

Step 2: Split the large file into 50k-line chunks

Split your main file into smaller, manageable chunks (each 50k lines):

split -l 50000 big_file.txt chunk_

This will create files like chunk_aa, chunk_ab, chunk_ac, etc.

Step 3: Process each chunk

  • Leave the first chunk (chunk_aa) as-is—it's our reference.
  • For every other chunk, replace lines 10-100 with the template:
for chunk in chunk_*; do
    if [ "$chunk" != "chunk_aa" ]; then
        # Keep first 9 lines, add template, then add lines 101+ from original chunk
        head -n 9 "$chunk" > "${chunk}_modified"
        cat template_lines.txt >> "${chunk}_modified"
        tail -n +101 "$chunk" >> "${chunk}_modified"
    fi
done

Step 4: Merge all modified chunks back into one file

cat chunk_aa chunk_*_modified > modified_big_file.txt

Cleanup (optional)

Delete the temporary chunks and template file once you verify the output is correct:

rm -f chunk_* template_lines.txt

Approach 2: Python Script (Flexible, No Intermediate Files)

If you prefer a single script that handles everything in one pass (no splitting into intermediate files), this Python solution streams the file line-by-line to avoid memory overload.

def replace_chunk_ranges(input_path, output_path, chunk_size=50000, start_line=10, end_line=100):
    template_lines = []
    current_line_num = 0

    with open(input_path, 'r') as infile, open(output_path, 'w') as outfile:
        for line in infile:
            current_line_num += 1
            # Calculate which chunk we're in (1-indexed)
            chunk_number = (current_line_num - 1) // chunk_size + 1
            # Calculate line position within the current chunk (1-indexed)
            line_in_chunk = current_line_num - (chunk_number - 1) * chunk_size

            if chunk_number == 1:
                # Save the template lines from the first chunk
                if start_line <= line_in_chunk <= end_line:
                    template_lines.append(line)
                # Write all lines from the first chunk as-is
                outfile.write(line)
            else:
                if line_in_chunk < start_line:
                    # Keep lines before the target range
                    outfile.write(line)
                elif line_in_chunk == start_line:
                    # Write the entire template once, skip the original lines in the range
                    outfile.writelines(template_lines)
                elif line_in_chunk <= end_line:
                    # Skip the original lines in the target range
                    continue
                else:
                    # Write lines after the target range
                    outfile.write(line)

# Run the function
replace_chunk_ranges("big_file.txt", "modified_big_file.txt")

Key Notes for Both Approaches

  • Backup first! Always make a copy of your original file before modifying it—accidents happen.
  • Edge cases: If the last chunk has fewer than 100 lines, both approaches handle it gracefully (the tail command will just output nothing, and the Python script will skip any lines beyond the chunk's end).
  • Performance: The command-line approach is generally faster for large files, while the Python script is more adaptable if you need to tweak the logic later.

内容的提问来源于stack exchange,提问作者nb20

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:49:06