Bash下针对大文件批量修剪指定行首字符的高效解决方案咨询
Hey there, let's fix this performance issue for your large file trimming task. Your current loop-based approach works, but as you've noticed, it's painfully slow for big files—every iteration rewrites the entire input1 file, and disk IO is the biggest bottleneck here. Repeating that hundreds or thousands of times kills performance.
Why Your Current Script Is Slow
Each time you run sed in the loop:
- It reads the full
input1file - Makes a single line edit
- Writes the entire file back to disk as
input1_new - Then you overwrite the original file with the new version
If your input2 has even a few hundred rules, that's hundreds of full file reads and writes—terrible for large datasets.
A Far More Efficient Approach: Use awk
We can handle this in one pass over both files, with only two reads (input1 and input2) and one write (output). Here's how:
First, we load all trimming rules into an associative array (mapping line numbers to the number of characters to trim). Then we process each line of input1 in order, applying the trim only if the current line has a matching rule.
Solution Script
Save this as trim_rules.awk:
BEGIN { # Load all trimming rules from input2 into an array while (getline rule_line < "input2") { # Split the line into individual rules (e.g., "2,1 1,2" becomes ["2,1", "1,2"]) num_rules = split(rule_line, rule_list, " ") for (i = 1; i <= num_rules; i++) { # Split each rule into (trim_count, line_number) split(rule_list[i], pair, ",") trim_length[pair[2]] = pair[1] # trim_length[line_num] = chars_to_trim } } close("input2") } # Process each line of input1 NR in trim_length { # Trim the first N characters (start printing from position N+1) print substr($0, trim_length[NR] + 1) next # Skip to next line, no need to print original } # If no rule applies, print the line as-is { print }
Run It Like This
awk -f trim_rules.awk input1 > output_file
Or use a one-liner (no separate script file):
awk ' BEGIN { while (getline < ARGV[2]) { n = split($0, rs, " ") for (i=1; i<=n; i++) { split(rs[i], p, ",") t[p[2]] = p[1] } } close(ARGV[2]) ARGV[2] = "" } NR in t { print substr($0, t[NR]+1); next } { print } ' input1 input2 > output_file
Key Advantages
- Minimal IO: Only reads
input1once,input2once, and writes the output once. No intermediate file overwrites. - Speed: Awk is optimized for text processing, and this approach avoids the overhead of spawning multiple
sed/awkprocesses in a loop. - Scalability: Works just as well for 10GB files as small ones, since storing the rules takes negligible memory compared to large files.
Edge Case Notes
- If
input2has duplicate rules for the same line number (e.g.,2,1 3,1), this script uses the last rule encountered. To trim the sum of duplicate counts, modify the array assignment totrim_length[pair[2]] += pair[1]. - Rules referencing lines that don't exist in
input1are simply ignored—no errors or unexpected behavior.
内容的提问来源于stack exchange,提问作者Wen Wang

