对比两文件查找非逆序重复词对并去重合并的高效实现
Efficient Solution to Remove Duplicate Consecutive Word Pairs and Merge Files
Your original approach has two key issues: it's inefficient (spawning multiple grep processes for each pair in File1) and incorrectly includes reversed pairs, which violates your requirement. Here's a streamlined, high-performance solution using awk that handles everything in one pass without temporary files.
Step-by-Step Explanation & Script
We'll use awk to:
- First scan File1 to build a set of all valid consecutive word pairs (ignoring reverses).
- Then process File2 to exclude any words that form a duplicate pair with the previous word.
- Finally, merge the original File1 content with the cleaned File2 content.
Save this as clean_and_merge.awk:
# Process File1 first: store its content and all valid word pairs NR == FNR { file1 = $0 word_count = split($0, words) for (i = 1; i < word_count; i++) { pair = words[i] " " words[i+1] duplicate_pairs[pair] = 1 } next } # Process File2 and output merged result { # Print original File1 content first print file1 prev_word = "" word_count = split($0, words) output = "" for (i = 1; i <= word_count; i++) { if (prev_word == "") { # First word of File2 is always kept output = words[i] prev_word = words[i] } else { current_pair = prev_word " " words[i] # Only keep the word if the pair isn't in our duplicate set if (!(current_pair in duplicate_pairs)) { output = output " " words[i] prev_word = words[i] } # Skip the word if the pair is duplicate; keep prev_word for next iteration } } # Print cleaned File2 content print output }
Run the Script
Execute it with your input files to get the merged result directly:
awk -f clean_and_merge.awk file1 file2 > merged_output.txt
Example Output
For your sample files:
- File1:
A B C D F A C G - File2:
A G C D A C D C F A
The script will generate merged_output.txt with:
A B C D F A C G A G C A D C F
(Duplicate pairs C D, A C, F A from File2 were removed as requested.)
Why This Is Better Than Your Original Approach
- Performance: Only reads each file once, no repeated
grepcalls or process spawning. This is critical for large files. - Accuracy: Strictly matches only forward pairs (reversed pairs like
D Caren't flagged as duplicates, per your requirement). - Simplicity: No temporary files—directly outputs the final merged result in one step.
内容的提问来源于stack exchange,提问作者Swimming bird
相关产品推荐
相关产品推荐

