You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用sed/grep替代循环高效统计近10分钟日志字符串出现次数

高效统计近10分钟日志中指定字符串的出现次数

问题背景

原本通过while循环每秒遍历一次日志文件统计指定字符串出现次数,这种方式需要反复打开、读取日志文件,效率极低,现需改用sed或grep命令替代循环提升效率。

原低效脚本

#!/bin/bash
log_file="logfile.log"
search_string="your_search_string"
COUNTER=0
current_time=$(date +%s)
start_time=$((current_time-600))
    while [ $start_time -le $current_time ]
    do
        
        date_time=$(date -d @$start_time +"%Y-%m-%d %H:%M:%S")
        count=$(grep "$date_time" "$log_file" | grep "$search_string" | wc -l)
            if [ $count -eq 1 ]; then
                COUNTER=$[$COUNTER +1]
                return
            fi
        echo "$date_time: $COUNTER occurrences"
        start_time=$((start_time+1))
    done
 echo "$COUNTER occurrences"

日志时间戳格式

日志中的时间戳格式为YYYY-MM-DD HH:MM:SS,mmm,示例如下:

INFO  [RMI TCP Connection(4)-10.103.5.24] 2022-11-29 15:21:01,552 Server.java:225 - Stop listening
WARN  [RMI TCP Connection(6)-10.103.5.24] 2022-11-29 15:21:07,948 StorageService.java:359 - Stopping gossip
INFO  [RMI TCP Connection(6)-10.103.5.24] 2022-11-29 15:21:07,949 Gossiper.java:1456 - Announcing shutdown

优化后的脚本

#!/bin/bash
log_file="logfile.log"
search_string="your_search_string"

# 计算近10分钟的起始时间字符串格式(精确到秒,忽略毫秒)
start_date=$(date -d "-10 minutes" +"%Y-%m-%d %H:%M:%S")

# 使用sed筛选出时间在起始时间之后的行,再用grep统计指定字符串出现次数
# 利用日志时间戳的字典序与时间顺序一致的特性,匹配时间戳前缀
count=$(sed -n "/^.*$start_date/p" "$log_file" | grep -c "$search_string")

echo "近10分钟内'$search_string'共出现$count次"

优化思路说明

  1. 一次性时间范围筛选:先计算出10分钟前的时间字符串(与日志时间戳前19位格式一致),借助日志时间戳的字典序和时间顺序一致的特性,用sed一次性筛选出所有符合时间范围的行,彻底避免循环遍历每秒时间的低效操作。
  2. 高效统计:用grep -c直接统计筛选后结果中指定字符串的出现次数,减少不必要的管道操作开销。
  3. 性能大幅提升:整个过程仅读取一次日志文件,相比原脚本的600次文件读取,效率提升显著。

进阶优化(适用于按时间递增的日志)

如果日志是按时间顺序递增写入的,可进一步缩小扫描范围,只读取起始时间对应的行之后的内容:

#!/bin/bash
log_file="logfile.log"
search_string="your_search_string"

start_date=$(date -d "-10 minutes" +"%Y-%m-%d %H:%M:%S")

# 找到第一个符合起始时间的行号
start_line=$(grep -n "^.*$start_date" "$log_file" | head -n1 | cut -d: -f1)

# 根据是否找到起始行,选择统计范围
if [ -n "$start_line" ]; then
    count=$(sed -n "${start_line},$p" "$log_file" | grep -c "$search_string")
else
    # 起始时间早于日志最早时间,统计整个文件
    count=$(grep -c "$search_string" "$log_file")
fi

echo "近10分钟内'$search_string'共出现$count次"

内容的提问来源于stack exchange,提问作者Junaid Khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 04:01:56