超2亿行大文本批量替换:用首块10-100行替换所有区块对应行
Got it, let's break this down—handling a 200M-line file sounds daunting, but by leaning into chunked processing and lightweight tools, we can get this done efficiently without blowing up your memory. Here are two solid approaches depending on your comfort with command-line tools vs. scripting:
Approach 1: Command-Line Toolchain (Fast, No Scripting)
This uses split, head, tail, and cat—all standard Unix/Linux tools that process files in streams (so no need to load the entire 200M-line file into memory).
Step 1: Extract the template lines from the first chunk
First, grab lines 10-100 from the first 50k-line chunk and save them to a temporary template file:
head -n 100 big_file.txt | tail -n 91 > template_lines.txt
(91 because 100-10+1 = 91 lines total)
Step 2: Split the large file into 50k-line chunks
Split your main file into smaller, manageable chunks (each 50k lines):
split -l 50000 big_file.txt chunk_
This will create files like chunk_aa, chunk_ab, chunk_ac, etc.
Step 3: Process each chunk
- Leave the first chunk (
chunk_aa) as-is—it's our reference. - For every other chunk, replace lines 10-100 with the template:
for chunk in chunk_*; do if [ "$chunk" != "chunk_aa" ]; then # Keep first 9 lines, add template, then add lines 101+ from original chunk head -n 9 "$chunk" > "${chunk}_modified" cat template_lines.txt >> "${chunk}_modified" tail -n +101 "$chunk" >> "${chunk}_modified" fi done
Step 4: Merge all modified chunks back into one file
cat chunk_aa chunk_*_modified > modified_big_file.txt
Cleanup (optional)
Delete the temporary chunks and template file once you verify the output is correct:
rm -f chunk_* template_lines.txt
Approach 2: Python Script (Flexible, No Intermediate Files)
If you prefer a single script that handles everything in one pass (no splitting into intermediate files), this Python solution streams the file line-by-line to avoid memory overload.
def replace_chunk_ranges(input_path, output_path, chunk_size=50000, start_line=10, end_line=100): template_lines = [] current_line_num = 0 with open(input_path, 'r') as infile, open(output_path, 'w') as outfile: for line in infile: current_line_num += 1 # Calculate which chunk we're in (1-indexed) chunk_number = (current_line_num - 1) // chunk_size + 1 # Calculate line position within the current chunk (1-indexed) line_in_chunk = current_line_num - (chunk_number - 1) * chunk_size if chunk_number == 1: # Save the template lines from the first chunk if start_line <= line_in_chunk <= end_line: template_lines.append(line) # Write all lines from the first chunk as-is outfile.write(line) else: if line_in_chunk < start_line: # Keep lines before the target range outfile.write(line) elif line_in_chunk == start_line: # Write the entire template once, skip the original lines in the range outfile.writelines(template_lines) elif line_in_chunk <= end_line: # Skip the original lines in the target range continue else: # Write lines after the target range outfile.write(line) # Run the function replace_chunk_ranges("big_file.txt", "modified_big_file.txt")
Key Notes for Both Approaches
- Backup first! Always make a copy of your original file before modifying it—accidents happen.
- Edge cases: If the last chunk has fewer than 100 lines, both approaches handle it gracefully (the
tailcommand will just output nothing, and the Python script will skip any lines beyond the chunk's end). - Performance: The command-line approach is generally faster for large files, while the Python script is more adaptable if you need to tweak the logic later.
内容的提问来源于stack exchange,提问作者nb20

