如何使用AWK仅提取匹配的目标ID字符串而非包含该ID的整行?
Got it, let's get that specific ID extracted properly! The issue with your original regex is that it doesn't account for the fixed-length segments in your target ID (8-4-4-4-12 characters separated by hyphens), which is why it was matching whole lines instead of just the ID. Here are a couple of AWK solutions tailored to your needs:
1. Extract Single ID Per Line
If each line has at most one occurrence of your target ID format, use this command. It leverages AWK's match() function to locate the ID, then substr() to pull out just the matching segment:
awk 'match($0, /[a-z0-9]{8}-[a-z0-9]{4}-[a-z0-9]{4}-[a-z0-9]{4}-[a-z0-9]{12}/) { print substr($0, RSTART, RLENGTH) }' your_input_file.txt
Breakdown:
- The regex
[a-z0-9]{8}-[a-z0-9]{4}-[a-z0-9]{4}-[a-z0-9]{4}-[a-z0-9]{12}exactly matches your ID's structure (8 characters, followed by three 4-char segments, then 12 characters, all separated by hyphens). match($0, regex)finds the position of the ID in the line, storing the starting index inRSTARTand the length of the match inRLENGTH.substr($0, RSTART, RLENGTH)extracts only the matched ID from the full line.
2. Extract Multiple IDs Per Line
If a single line might contain multiple instances of this ID format, use a loop to keep extracting until there are no more matches left:
awk '{ current_line = $0 while (match(current_line, /[a-z0-9]{8}-[a-z0-9]{4}-[a-z0-9]{4}-[a-z0-9]{4}-[a-z0-9]{12}/)) { print substr(current_line, RSTART, RLENGTH) current_line = substr(current_line, RSTART + RLENGTH) } }' your_input_file.txt
Example Usage:
Suppose your input line looks like this:
User action: ID=549d40r0-1e1b-01v8-0d72-0a8f2c680100, linked ID=xyz789ab-123c-456d-7890-abcdef123456
Running the first command would output:
549d40r0-1e1b-01v8-0d72-0a8f2c680100
Running the second command would output both IDs on separate lines:
549d40r0-1e1b-01v8-0d72-0a8f2c680100 xyz789ab-123c-456d-7890-abcdef123456
内容的提问来源于stack exchange,提问作者user3834663

