You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Linux下如何用sed/grep输出正则命名捕获组的匹配名称?

Solution Using Awk + Sed (No PHP Required)

Great question! Handling large 200-500MB text files with basic Linux tools while mapping matches to regex named capture groups is totally feasible. While sed and grep alone lack the ability to easily link matched patterns to their group names/numbers, combining them with awk (a standard, efficient Linux utility) lets you build a pattern-to-group mapping and process the file efficiently.

Step 1: Understand the Approach

We’ll break this into two core parts:

  1. Extract the named capture groups from your regex to create a mapping of pattern → group name/number.
  2. Use this mapping to scan your input file, match words, and output the corresponding group identifiers in your desired format.

Step 2: Example Implementation

Let’s use your sample regex: (?P<id1>aba)|(?P<id2>cde)|(?P<id3>esa)|(?P<id4>fav) and input file input.txt with content aba cec fav gex hur.

Variant 1: Output Group Names (Space-Separated)

# Define your regex
REGEX='(?P<id1>aba)|(?P<id2>cde)|(?P<id3>esa)|(?P<id4>fav)'

# Generate a pattern-to-group-name mapping for awk
MAP=$(echo "$REGEX" | sed -E 's/\(\?P<([^>]+)>([^)]+)\)/pattern["\2"]="\1";/g; s/\|//g')

# Process the input file
awk -v map="$MAP" '
BEGIN {
    # Load the mapping into awk variables
    eval(map)
}
{
    result = ""
    # Check each word in the line
    for (i=1; i<=NF; i++) {
        if ($i in pattern) {
            # Build the result string with spaces
            result = (result != "") ? result " " pattern[$i] : pattern[$i]
        }
    }
    # Print only if there are matches
    if (result != "") print result
}' input.txt

This will output: id1 id4

Variant 2: Output Group Names (Semicolon-Separated)

Just modify the separator in the awk logic:

awk -v map="$MAP" '
BEGIN { eval(map) }
{
    result = ""
    for (i=1; i<=NF; i++) {
        if ($i in pattern) {
            result = (result != "") ? result ";" pattern[$i] : pattern[$i]
        }
    }
    if (result != "") print result
}' input.txt

Output: id1;id4

Variant 3: Output Group Numbers (Space-Separated)

Adjust the mapping to extract just the numeric part of the group names:

MAP=$(echo "$REGEX" | sed -E 's/\(\?P<id([0-9]+)>([^)]+)\)/pattern["\2"]="\1";/g; s/\|//g')

# Reuse the same awk command as Variant 1
awk -v map="$MAP" '
BEGIN { eval(map) }
{
    result = ""
    for (i=1; i<=NF; i++) {
        if ($i in pattern) {
            result = (result != "") ? result " " pattern[$i] : pattern[$i]
        }
    }
    if (result != "") print result
}' input.txt

Output: 1 4

Variant 4: Output Group Numbers (Semicolon-Separated)

Combine the number extraction with semicolon separation:

awk -v map="$MAP" '
BEGIN { eval(map) }
{
    result = ""
    for (i=1; i<=NF; i++) {
        if ($i in pattern) {
            result = (result != "") ? result ";" pattern[$i] : pattern[$i]
        }
    }
    if (result != "") print result
}' input.txt

Output: 1;4

Key Notes for Large Files

  • Efficiency: awk is optimized for large text processing and will handle 200-500MB files smoothly, as it processes lines incrementally without loading the entire file into memory.
  • Regex Flexibility: If your named groups don’t follow the idN format, adjust the sed command to extract group names or assign sequential numbers. For example, to assign numbers based on group order:
    MAP=$(echo "$REGEX" | sed -E 's/\(\?P<([^>]+)>([^)]+)\)/&\n/g' | grep -E '\(\?P<' | nl -v1 | sed -E 's/^[0-9]+[[:space:]]+\(\?P<([^>]+)>([^)]+)\)/pattern["\2"]="\1"; num_pattern["\2"]="\1";/g' | tr -d '\n')
    
  • Special Characters: If your regex patterns contain parentheses or other special characters, tweak the sed matching to avoid breaking the mapping extraction (e.g., use more precise regex to capture patterns).

内容的提问来源于stack exchange,提问作者AmpliDude

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 09:17:41