使用Grep提取带计数的唯一报错失败,求助去除后缀数字/主机名
Got it, let's sort out this issue! The problem here is that your current command captures the full CCS ERROR line—including those dynamic trailing values like numbers and hostnames—which makes uniq treat identical error messages as unique entries. We just need to tweak the regex to extract only the static error description part, stripping off the variable suffixes.
Modified Command
Here's an updated version of your script that targets the core error message (adjust the regex if your log format varies slightly):
grep -n 'CCS ERROR.*' $dir/*.csv* | grep -oE 'CCS ERROR:[^,]+' | sort | uniq -c | sort -nr >> /path/output.txt
How It Works
Let's break down the key change:
grep -oE 'CCS ERROR:[^,]+': The[^,]+regex matches any sequence of characters that aren't commas. In your example log line (martpay:signal2:auth:CCS ERROR: PayLoad are not matching,51767703062,sXXXXpgsusapp620), this extracts justCCS ERROR: PayLoad are not matching—cutting off the comma-separated dynamic values entirely.
Adapting to Other Log Formats
If your suffixes aren't comma-separated, adjust the regex to match your specific pattern:
- Suffix starts with a number: Use PCRE mode (
-P) to stop at the first space followed by a number:grep -n 'CCS ERROR.*' $dir/*.csv* | grep -oP 'CCS ERROR:.*?(?=\s+\d)' | sort | uniq -c | sort -nr >> /path/output.txt - Suffix is a hostname starting with
s: Stop at the first space followed by ans-prefixed hostname:grep -n 'CCS ERROR.*' $dir/*.csv* | grep -oP 'CCS ERROR:.*?(?=\s+s[a-zA-Z0-9]+)' | sort | uniq -c | sort -nr >> /path/output.txt
Test this with your logs first to make sure it captures exactly the static error text you need—tweak the regex boundaries if you see any missing or extra content!
内容的提问来源于stack exchange,提问作者user3489078

