日志提取脚本优化:需包含指定时间范围内的无时间戳日志行
日志提取脚本优化:包含时间范围内的无时间戳关联日志行
问题描述
编写的Bash脚本仅能提取指定时间范围内带有时间戳的日志行,遗漏了处于该时间范围内但无时间戳的关联日志行(如示例中的a、b、c行)。需要优化脚本以覆盖这类无时间戳行,且无法修改现有时间戳格式。
日志示例
[17:02:12:161][01-03-2024]some log info here: step1 step2 step3 [17:02:12:163][01-03-2024]some log here a b c [17:02:12:185][01-03-2024]Time taken : 11
指定时间范围
- 开始时间戳:
[17:02:12:161][01-03-2024] - 结束时间戳:
[17:02:12:163][01-03-2024]
原脚本问题
原脚本的awk逻辑仅打印本身带有符合条件时间戳的行,未跟踪当前是否处于目标时间区间的状态,导致时间戳行后续的无时间戳关联行被遗漏。
优化后的脚本
#!/bin/bash # Function to check the timestamp format timestamp_pattern_checker() { local input_pattern="^\[([0-9]{2}:[0-9]{2}:[0-9]{2}:[0-9]{3})\]\[([0-9]{2}-[0-9]{2}-[0-9]{4})\]$" if [[ ! $1 =~ $input_pattern ]]; then echo "Invalid Timestamp pattern" echo "Timestamps should be in the format '[HH:MM:SS:SSS][DD-MM-YYYY]' or HH" exit 1 fi } # Function to convert hours to timestamp format convert_hours_to_timestamp() { local hour=$1 printf "[%02d:00:00:000]" "$hour" } # Check if correct number of arguments are passed if [ "$#" -ne 3 ]; then echo "Usage: $0 <log_file_name> <start_timestamp> <end_timestamp>" echo "Timestamps should be in the format '[HH:MM:SS:SSS][DD-MM-YYYY]' or as integers representing hours (00 to 23)" exit 1 fi log_file=$1 start_input=$2 end_input=$3 # Check if log file exists if [ ! -f "$log_file" ]; then echo "File not found: $log_file" exit 1 fi # Determine if inputs are hours or full timestamps if [[ $start_input =~ ^[0-9]{2}$ && $end_input =~ ^[0-9]{2}$ ]]; then if [[ $start_input -ge 0 && start_input -le 23 && $end_input -ge 0 && $end_input -le 23 ]]; then if [[ $start_input -le $end_input ]]; then start_timestamp=$(convert_hours_to_timestamp "$start_input") end_timestamp=$(convert_hours_to_timestamp "$end_input") # Extract unique dates from log file unique_dates=$(awk -F'[][]' '{print $4}' "$log_file" | sort | uniq) start_timestamps=() end_timestamps=() for date in $unique_dates; do start_timestamps+=("${start_timestamp}[$date]") end_timestamps+=("${end_timestamp}[$date]") done else echo "Error: start hour must be less than or equal to end hour." exit 1 fi else echo "Error: Hours must be between 00 and 23." exit 1 fi else start_timestamp=$start_input end_timestamp=$end_input timestamp_pattern_checker "$start_timestamp" timestamp_pattern_checker "$end_timestamp" start_timestamps=("$start_timestamp") end_timestamps=("$end_timestamp") fi awk -v starts="${start_timestamps[*]}" -v ends="${end_timestamps[*]}" ' function parsedate(date) { split(date, a, /[]:[-]+/) # 处理带时间戳的行,返回ISO格式时间;无时间戳的行返回空 if (a[1] != "") { return a[6] "-" a[7] "-" a[8] "T" a[1] ":" a[2] ":" a[3] "." a[4] } else { return "" } } BEGIN { split(starts, start_arr, " ") split(ends, end_arr, " ") for (i in start_arr) { st[i] = parsedate(start_arr[i]) et[i] = parsedate(end_arr[i]) } in_range = 0 # 标记是否处于目标时间范围内 log_count = 0 } { current_time = parsedate($0) # 如果当前行是带时间戳的行,判断是否在目标区间内 if (current_time != "") { in_range = 0 for (i in st) { if (current_time >= st[i] && current_time <= et[i]) { in_range = 1 break } # 额外判断:如果当前时间戳超过所有结束时间,直接关闭输出 if (current_time > et[i]) { in_range = 0 } } } # 如果处于目标区间内,打印当前行 if (in_range) { print $0 log_count = 1 } } END { if (log_count == 0) { print "Nothing was logged at this given time frame" } }' "$log_file"
关键优化点
- 新增
in_range变量,持续跟踪当前是否处于目标时间区间内 - 修改
parsedate函数,对无时间戳的行返回空值,明确区分带时间戳和不带时间戳的行 - 调整核心逻辑:
- 遇到带时间戳的行时,重新评估是否处于目标区间,更新
in_range状态 - 无论当前行是否带时间戳,只要
in_range为1(处于区间内),就打印该行 - 当时间戳超过所有结束时间时,自动关闭输出状态,避免后续无关行被误输出
- 遇到带时间戳的行时,重新评估是否处于目标区间,更新
内容的提问来源于stack exchange,提问作者Arav Stark
相关产品推荐
相关产品推荐

