如何用Linux基础工具统计日志中按时间戳分组的错误数?
统计日志中的错误事件数(基于Linux基础工具)
需要统计日志中的错误事件数,规则是每个时间戳后跟随的多行错误属于同一个事件,而非按错误行数统计。
日志示例
2023-09-25 11:11:02.4594 2023-09-25 11:12:02.9254 2023-09-25 11:13:03.0777 2023-09-25 11:14:02.4585 fetchmail: WARNING: imap.xy.z.ch configuration invalid, you normally need port 993/service imaps for --ssl. fetchmail: OpenSSL reported: error:1408F10B:SSL routines:ssl3_get_record:wrong version number fetchmail: imap.xy.z.ch : SSL connection failed. fetchmail: socket error while fetching from user@xy.ch@imap.xy.z.ch 2023-09-25 11:15:02.6085 fetchmail: WARNING: imap.xy.z.ch configuration invalid, you normally need port 993/service imaps for --ssl. fetchmail: OpenSSL reported: error:1408F10B:SSL routines:ssl3_get_record:wrong version number fetchmail: imap.xy.z.ch : SSL connection failed. fetchmail: socket error while fetching from user@xy.ch@imap.xy.z.ch 2023-09-25 11:16:02.2786 fetchmail: WARNING: imap.xy.z.ch configuration invalid, you normally need port 993/service imaps for --ssl. fetchmail: OpenSSL reported: error:1408F10B:SSL routines:ssl3_get_record:wrong version number fetchmail: imap.xy.z.ch : SSL connection failed. fetchmail: socket error while fetching from user@xy.ch@imap.xy.z.ch 2023-09-25 11:17:02.5138 fetchmail:/somewhere/my.fetchmail:13: syntax error at user 2023-09-25 11:18:02.7093 2023-09-25 11:19:02.9584 2023-09-25 11:20:02.5294
当前问题与现有脚本
直接统计所有非时间戳行得到13个"错误",但实际应统计为4个错误事件。我自己编写了一段bash脚本可以得到正确结果:
#!/bin/bash if [ "$1" == "" ]; then echo "Log filename required - stopping" exit 1 fi logFile=$1 found=0 shortDate=$(date +%Y-%m-%d) echo "** Beginn..." while read line do curr=${line:0:10} if [ "$last" == "" ]; then last=${line:0:10} fi if [ "$last" != "$curr" ]; then if [ "$curr" != "$shortDate" ]; then let found+=1 fi fi last=$curr done < $logFile echo $found
简洁解决方案
方法1:使用awk直接统计
awk '/^[0-9]{4}-[0-9]{2}-[0-9]{2}/ { if (prev_line && prev_line !~ /^[0-9]{4}-[0-9]{2}-[0-9]{2}/) count++ } {prev_line=$0} END { if (prev_line !~ /^[0-9]{4}-[0-9]{2}-[0-9]{2}/) count++ print count }' logfile.txt
逻辑:遍历日志时记录上一行内容;遇到新时间戳行时,若上一行是错误行(非时间戳)则计数加1;最后检查日志末尾是否以错误行结束,若是则补充计数。
方法2:grep组合命令
grep -B1 -v "^[0-9]\{4\}-[0-9]\{2\}-[0-9]\{2\}" logfile.txt | grep "^[0-9]\{4\}-[0-9]\{2\}-[0-9]\{2\}" | wc -l
逻辑:
grep -B1 -v "^[0-9]{4}-[0-9]{2}-[0-9]{2}":找出所有非时间戳的错误行,并显示它们的前一行(触发事件的时间戳)- 过滤出这些时间戳行,统计行数即为错误事件数
方法3:sed命令实现
sed -n '/^[0-9]\{4\}-[0-9]\{2\}-[0-9]\{2\}/{n;/^fetchmail:/{x;p;x}}' logfile.txt | wc -l
逻辑:找到时间戳行后,检查下一行是否为fetchmail:开头的错误行;若是则打印该时间戳行,最后统计打印的行数。
内容的提问来源于stack exchange,提问作者Pascal Kuenzli
相关产品推荐
相关产品推荐

