You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Bash脚本日志解析:捕获错误及续行并去重的实现问题

日志文件ERROR条目提取与去重方案

我是Bash脚本新手,需要解析日志文件并将目标信息输出到TXT文件,具体需求如下:

  • 仅提取ERROR条目及其后续非日期开头的续行(如堆栈跟踪)
  • 输出时去除日期、时间、线程信息
  • 重复的错误内容需去重

当前用grep、cut、sort、uniq实现了基础功能,但无法捕获错误的续行,尝试grep加条件逻辑未成功,正打算改用awk解决。


日志示例

2022-1-3 14:00:00 ERROR THREAD234 - error info here 
2022-1-4 02:00:00 WARNI THREAD235 - warning info here 
additional warning info here, sometimes includes word error, but i do not want to capture this additional line as it is a warning
2022-2-3 01:00:00 ERROR THREAD333 - error info here2 
error info continued, sometimes there are multiple lines to an error and they do not all include the word error. however, these additional lines to do not include date/times. these are typically stack traces. 
2023-3-4 11:00:00 INFO0 THREAD333 - info here
2022-2-5 01:00:00 ERROR THREAD333 - error info here3
2022-2-6 06:00:00 ERROR THREAD333 - error info here3

期望输出

1 ERROR - error info here
1 ERROR - error info here2
error info continued, sometimes there are multiple lines to an error and they do not all include the word error. however, these additional lines to do not include date/times. these are typically stack traces.
2 ERROR - error info here3

当前输出

1 ERROR - error info here
1 ERROR - error info here2
2 ERROR - error info here3

现有脚本

#!/bin/bash

read -p "File path to log, no spaces: " file
outputFile=Desktop/errorOutput.txt
error=$(grep ERROR $file | cut -b 25-32,47-1000 | sort | uniq -c)
touch $outputFile
echo "$error" > $outputFile
cat $outputFile

解决方案:使用awk实现完整需求

以下awk脚本可捕获ERROR条目及其续行,去除无关信息并完成去重计数:

#!/bin/bash

read -p "File path to log, no spaces: " file
outputFile=~/Desktop/errorOutput.txt

awk '
# 匹配以日期开头的ERROR行
/^[0-9]{4}-[0-9]{1,2}-[0-9]{1,2}/ && /ERROR/ {
    # 移除日期、时间、线程前缀,保留ERROR及后续内容
    sub(/^[0-9]{4}-[0-9]{1,2}-[0-9]{1,2} [0-9]{2}:[0-9]{2}:[0-9]{2} ERROR [^ ]+ - /, "ERROR - ", $0)
    current_error = $0
    # 初始化当前错误的完整内容
    error_lines[current_error] = current_error
    in_error = 1
    next
}
# 处于ERROR续行状态,且当前行不以日期开头
in_error && !/^[0-9]{4}-[0-9]{1,2}-[0-9]{1,2}/ {
    # 将续行追加到当前错误内容中
    error_lines[current_error] = error_lines[current_error] "\n" $0
    next
}
# 遇到非ERROR的日期行,退出续行捕获状态
/^[0-9]{4}-[0-9]{1,2}-[0-9]{1,2}/ {
    in_error = 0
}
# 处理完成后输出去重并计数的结果
END {
    # 统计每个完整错误的出现次数
    for (err in error_lines) {
        count[err]++
    }
    # 按内容排序输出(与原有脚本逻辑一致)
    PROCINFO["sorted_in"] = "@ind_str_asc"
    for (err in count) {
        print count[err], err
    }
}
' "$file" > "$outputFile"

cat "$outputFile"

脚本逻辑说明

  1. 识别ERROR起始行:匹配以日期开头且包含ERROR的行,去除无关前缀后记录当前错误内容,标记进入续行捕获状态。
  2. 捕获续行内容:若处于续行状态且当前行不以日期开头,将该行追加到对应错误的完整内容中。
  3. 退出续行状态:遇到其他日期开头的行时,结束当前错误的续行捕获。
  4. 去重与计数:用数组存储包含续行的完整错误内容,统计重复次数后按内容排序输出。

内容的提问来源于stack exchange,提问作者Moe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 06:12:35