Linux下以特殊字符为起止提取关键词前后字符串的日志处理需求
Since your warning logs follow a consistent structure (Warning: Variants 'XXX' and 'YYY' have the same position.), we can leverage Linux's built-in text-processing tools to extract the variant IDs using single quotes (') as the delimiters. Here are a few straightforward methods:
Method 1: Use awk (Great for structured text)
awk makes it easy to split lines by a custom delimiter and pick specific fields directly. We'll use single quotes as the separator:
awk -F"'" '{print $2, $4}' your_log_file.log
-F"'": Sets the field separator to a single quote$2: Grabs the first variant ID (between the first and second single quote)$4: Grabs the second variant ID (between the third and fourth single quote)- Add
> extracted_variants.txtat the end to save results to a file instead of printing to the terminal.
Method 2: Use sed (Pattern-based replacement)
sed can match the entire warning line and replace it with just the variant IDs using capture groups:
sed "s/Warning: Variants '\(.*\)' and '\(.*\)' have the same position./\1 \2/" your_log_file.log
- The regex
'\(.*\)'captures the text inside each pair of single quotes \1and\2refer to the first and second captured groups, so we replace the whole line with just those two IDs.
Method 3: Use grep (GNU grep with Perl regex)
If you have GNU grep (standard on most Linux distros), you can use Perl-compatible regex to directly extract the IDs:
grep -Po "Variants '\K[^']+' and '\K[^']+" your_log_file.log
-P: Enables Perl regex mode\K: Tells grep to "forget" the text matched before this point, so we only get the part after it[^']+: Matches any character except a single quote (exactly the variant ID)
All these methods will output each pair of variant IDs on a single line, matching the structure of your input logs.
内容的提问来源于stack exchange,提问作者Yatrosin

