You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

对比两文件查找非逆序重复词对并去重合并的高效实现

Efficient Solution to Remove Duplicate Consecutive Word Pairs and Merge Files

Your original approach has two key issues: it's inefficient (spawning multiple grep processes for each pair in File1) and incorrectly includes reversed pairs, which violates your requirement. Here's a streamlined, high-performance solution using awk that handles everything in one pass without temporary files.

Step-by-Step Explanation & Script

We'll use awk to:

  1. First scan File1 to build a set of all valid consecutive word pairs (ignoring reverses).
  2. Then process File2 to exclude any words that form a duplicate pair with the previous word.
  3. Finally, merge the original File1 content with the cleaned File2 content.

Save this as clean_and_merge.awk:

# Process File1 first: store its content and all valid word pairs
NR == FNR {
    file1 = $0
    word_count = split($0, words)
    for (i = 1; i < word_count; i++) {
        pair = words[i] " " words[i+1]
        duplicate_pairs[pair] = 1
    }
    next
}

# Process File2 and output merged result
{
    # Print original File1 content first
    print file1
    
    prev_word = ""
    word_count = split($0, words)
    output = ""
    
    for (i = 1; i <= word_count; i++) {
        if (prev_word == "") {
            # First word of File2 is always kept
            output = words[i]
            prev_word = words[i]
        } else {
            current_pair = prev_word " " words[i]
            # Only keep the word if the pair isn't in our duplicate set
            if (!(current_pair in duplicate_pairs)) {
                output = output " " words[i]
                prev_word = words[i]
            }
            # Skip the word if the pair is duplicate; keep prev_word for next iteration
        }
    }
    
    # Print cleaned File2 content
    print output
}

Run the Script

Execute it with your input files to get the merged result directly:

awk -f clean_and_merge.awk file1 file2 > merged_output.txt

Example Output

For your sample files:

  • File1: A B C D F A C G
  • File2: A G C D A C D C F A

The script will generate merged_output.txt with:

A B C D F A C G
A G C A D C F

(Duplicate pairs C D, A C, F A from File2 were removed as requested.)

Why This Is Better Than Your Original Approach

  • Performance: Only reads each file once, no repeated grep calls or process spawning. This is critical for large files.
  • Accuracy: Strictly matches only forward pairs (reversed pairs like D C aren't flagged as duplicates, per your requirement).
  • Simplicity: No temporary files—directly outputs the final merged result in one step.

内容的提问来源于stack exchange,提问作者Swimming bird

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:46:20