You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWK数组精准搜索求助:实现逐字符全匹配所有组合元素

Got it, let's tackle this precise character-position matching task with AWK. The core requirement here is to scan every possible starting character position in sample.txt for exact matches against the 6 elements from your combinations file—no block-based searching, just sliding-window character-level checks. And we want the most compact results, which I'll cover in two practical interpretations: all unique exact matches, or the shortest match at each position (since "compact" often implies minimal length).


Basic Script: All Unique Exact Matches

This script loads your target strings, scans every character position in sample.txt, and outputs all unique matches without duplicates.

# Load target strings from combinations file first
BEGIN {
    # Read each line from combinations, store string and its length for efficiency
    while ((getline t < "combinations") > 0) {
        # Skip blank lines to avoid unwanted matches
        if (t ~ /^[[:space:]]*$/) continue
        targets[t] = length(t)
    }
    close("combinations")
}

# Process each line in sample.txt
{
    current_line = $0
    line_length = length(current_line)
    # Track seen matches to avoid duplicate outputs
    delete seen_matches

    # Loop through every possible starting character position (1-based in AWK)
    for (start_pos = 1; start_pos <= line_length; start_pos++) {
        # Check each target string against the substring starting at start_pos
        for (target in targets) {
            target_len = targets[target]
            # Skip if target is longer than remaining characters in the line
            if (start_pos + target_len - 1 > line_length) continue
            
            # Extract the substring to compare
            candidate = substr(current_line, start_pos, target_len)
            
            # Exact match found?
            if (candidate == target) {
                # Create a unique key to avoid duplicate outputs
                match_key = NR "," start_pos "," target
                if (!(match_key in seen_matches)) {
                    printf "Line %d, Start at position %d: Matched '%s'\n", NR, start_pos, target
                    seen_matches[match_key] = 1
                }
            }
        }
    }
}

How to Use:

Save the script as char_match.awk, then run it with:

awk -f char_match.awk sample.txt

What This Does:

  • The BEGIN block preloads all non-blank lines from combinations into an array, storing each target and its length (avoids recalculating length repeatedly).
  • For every line in sample.txt, we loop through every character position as the starting point.
  • We check if any target fits exactly starting at that spot, then use a tracking array to skip duplicate matches for the same line, position, and target.

Compact Version: Shortest Match at Each Position

If "most compact" means prioritizing the shortest matching string at each starting position (since shorter strings are more compact), use this modified script:

BEGIN {
    while ((getline t < "combinations") > 0) {
        if (t ~ /^[[:space:]]*$/) continue
        targets[t] = length(t)
    }
    close("combinations")
}

{
    current_line = $0
    line_length = length(current_line)
    # Track the shortest match for each starting position
    delete shortest_matches

    for (start_pos = 1; start_pos <= line_length; start_pos++) {
        shortest_len = 999999  # Initialize with a large placeholder
        shortest_match = ""

        for (target in targets) {
            target_len = targets[target]
            if (start_pos + target_len - 1 > line_length) continue
            
            candidate = substr(current_line, start_pos, target_len)
            if (candidate == target) {
                # Update if this target is shorter than the current shortest
                if (target_len < shortest_len) {
                    shortest_len = target_len
                    shortest_match = target
                }
            }
        }

        # Print the shortest match if found
        if (shortest_match != "") {
            printf "Line %d, Start at position %d: Compact match '%s' (length %d)\n", NR, start_pos, shortest_match, shortest_len
        }
    }
}

Key Notes:

  • Both scripts use exact string comparison (not regular expressions), so special characters in your target strings won't cause unexpected behavior—perfect for your "精准搜索" requirement.
  • AWK uses 1-based indexing for strings, so position 1 is the first character of the line. If you need 0-based indexing, just subtract 1 from start_pos in the output print statements.
  • Duplicate lines in combinations are automatically deduplicated in both scripts, since we use the target string as the array key.

内容的提问来源于stack exchange,提问作者user289944

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:00:25