Bash脚本日志解析:捕获错误及续行并去重的实现问题
日志文件ERROR条目提取与去重方案
我是Bash脚本新手,需要解析日志文件并将目标信息输出到TXT文件,具体需求如下:
- 仅提取ERROR条目及其后续非日期开头的续行(如堆栈跟踪)
- 输出时去除日期、时间、线程信息
- 重复的错误内容需去重
当前用grep、cut、sort、uniq实现了基础功能,但无法捕获错误的续行,尝试grep加条件逻辑未成功,正打算改用awk解决。
日志示例
2022-1-3 14:00:00 ERROR THREAD234 - error info here 2022-1-4 02:00:00 WARNI THREAD235 - warning info here additional warning info here, sometimes includes word error, but i do not want to capture this additional line as it is a warning 2022-2-3 01:00:00 ERROR THREAD333 - error info here2 error info continued, sometimes there are multiple lines to an error and they do not all include the word error. however, these additional lines to do not include date/times. these are typically stack traces. 2023-3-4 11:00:00 INFO0 THREAD333 - info here 2022-2-5 01:00:00 ERROR THREAD333 - error info here3 2022-2-6 06:00:00 ERROR THREAD333 - error info here3
期望输出
1 ERROR - error info here 1 ERROR - error info here2 error info continued, sometimes there are multiple lines to an error and they do not all include the word error. however, these additional lines to do not include date/times. these are typically stack traces. 2 ERROR - error info here3
当前输出
1 ERROR - error info here 1 ERROR - error info here2 2 ERROR - error info here3
现有脚本
#!/bin/bash read -p "File path to log, no spaces: " file outputFile=Desktop/errorOutput.txt error=$(grep ERROR $file | cut -b 25-32,47-1000 | sort | uniq -c) touch $outputFile echo "$error" > $outputFile cat $outputFile
解决方案:使用awk实现完整需求
以下awk脚本可捕获ERROR条目及其续行,去除无关信息并完成去重计数:
#!/bin/bash read -p "File path to log, no spaces: " file outputFile=~/Desktop/errorOutput.txt awk ' # 匹配以日期开头的ERROR行 /^[0-9]{4}-[0-9]{1,2}-[0-9]{1,2}/ && /ERROR/ { # 移除日期、时间、线程前缀,保留ERROR及后续内容 sub(/^[0-9]{4}-[0-9]{1,2}-[0-9]{1,2} [0-9]{2}:[0-9]{2}:[0-9]{2} ERROR [^ ]+ - /, "ERROR - ", $0) current_error = $0 # 初始化当前错误的完整内容 error_lines[current_error] = current_error in_error = 1 next } # 处于ERROR续行状态,且当前行不以日期开头 in_error && !/^[0-9]{4}-[0-9]{1,2}-[0-9]{1,2}/ { # 将续行追加到当前错误内容中 error_lines[current_error] = error_lines[current_error] "\n" $0 next } # 遇到非ERROR的日期行,退出续行捕获状态 /^[0-9]{4}-[0-9]{1,2}-[0-9]{1,2}/ { in_error = 0 } # 处理完成后输出去重并计数的结果 END { # 统计每个完整错误的出现次数 for (err in error_lines) { count[err]++ } # 按内容排序输出(与原有脚本逻辑一致) PROCINFO["sorted_in"] = "@ind_str_asc" for (err in count) { print count[err], err } } ' "$file" > "$outputFile" cat "$outputFile"
脚本逻辑说明
- 识别ERROR起始行:匹配以日期开头且包含ERROR的行,去除无关前缀后记录当前错误内容,标记进入续行捕获状态。
- 捕获续行内容:若处于续行状态且当前行不以日期开头,将该行追加到对应错误的完整内容中。
- 退出续行状态:遇到其他日期开头的行时,结束当前错误的续行捕获。
- 去重与计数:用数组存储包含续行的完整错误内容,统计重复次数后按内容排序输出。
内容的提问来源于stack exchange,提问作者Moe
相关产品推荐
相关产品推荐

