大文件正则匹配替换优化:将头行10位编号批量追加至数据行
Oh, I feel your pain—repeating a regex replacement 200 times on huge files is such a waste of time and system resources. Let's fix this properly with solutions that process your files once instead of 200 times, which will cut down runtime dramatically.
First, Why Your Current Method Is Slow
Your current regex only handles one data row per header per replacement pass. Running it 200 times means you're scanning the entire file 200 times—no wonder it's taking forever. We need to process each line exactly once, keeping track of the current header ID as we go.
Solution 1: Use Awk (Fastest for Command-Line Users)
Awk is built for exactly this kind of line-by-line, block-based text processing, and it's lightning-fast even on massive files. Here's how to do it:
Run this command in your terminal (works on Linux, macOS, and Windows with WSL/Git Bash):
awk '/^[0-9]{10}\t/ { hdr = substr($0, 1, 10) # Grab the 10-digit header ID print $0 # Print the original header line next # Skip to the next line } { print hdr "\t" $0 # Prepend the stored header ID to data lines }' input.txt > output.txt
How It Works:
- When it hits a line starting with 10 digits + tab, it saves the 10-digit ID and prints the header line as-is.
- For every other line (your data rows), it prints the saved header ID, a tab, then the original data line.
- This processes the file once—no repeated scans.
Solution 2: Python Script (Cross-Platform, Flexible)
If you prefer a cross-platform solution you can tweak easily, use a Python script. It reads the file line by line (so it doesn't load the entire file into memory, critical for large files):
with open("input.txt", "r", encoding="utf-8") as infile, open("output.txt", "w", encoding="utf-8") as outfile: current_header = "" for line in infile: stripped_line = line.rstrip("\r\n") # Normalize line endings # Check if this is a header line (10 digits + tab at start) if len(stripped_line) >= 11 and stripped_line[:10].isdigit() and stripped_line[10] == "\t": current_header = stripped_line[:10] outfile.write(line) # Write original header line else: # Prepend header ID to data line outfile.write(f"{current_header}\t{stripped_line}\n")
Why This Is Better:
- Handles any number of data rows per header (no hardcoded 200 limit).
- Works on Windows, macOS, and Linux without extra tools (just Python installed).
- Uses minimal memory, even for multi-GB files.
Solution 3: EmEditor Script Replacement (Stay in Your Editor)
If you want to stick with EmEditor instead of switching tools, use its built-in JavaScript script replacement to process the file in one pass:
- Open your file in EmEditor.
- Press
Ctrl+Hto open the Replace window. - Click the Advanced button, then check the Use Script box.
- Paste this script into the text box:
var currentHeader = ""; function replaceLine(line) { // Check if line starts with 10 digits + tab if (/^\d{10}\t/.test(line)) { currentHeader = line.substring(0, 10); return line; } else { return currentHeader + "\t" + line; } }
- Click All Replace—EmEditor will process every line once, appending the correct header ID to each data row.
All These Solutions Do One Thing Right:
They only scan your file once. No more repeated regex passes, no more waiting hours for large files to process. Pick the one that fits your workflow best!
内容的提问来源于stack exchange,提问作者vunsb

