如何提取并排序AV日志中扫描耗时与扫描文件数的最大值
问题描述
我有一份AV日志文件,每条进程条目包含Name、Path、Total files scanned、Scan time等信息,文件里有数百条这类记录。我需要对Total files scanned和Scan time字段分别排序,找出数值最大的记录,同时关联对应的Process id(比如显示“Process id: 86, Scan time (ns): 12761174”这样的格式,按数值从大到小排列)。
之前用grep命令只能得到单独的数字排序列表,没法关联进程ID,尝试的命令如下:
grep -Eo | grep 'Scan time (ns)' '[0-9]+' file | sort
得到的结果只有扫描时间和文件名,没有进程ID:
file:Scan time (ns): "9391986" file:Scan time (ns): "9532119" file:Scan time (ns): "9730650" file:Scan time (ns): "9743828" file:Scan time (ns): "9793469" file:Scan time (ns): "9911768"
期望的输出效果(按扫描时间降序排列):
Process id 19, Scan time (ns): "178764794932" Process id 30, Scan time (ns): "90257651" Process id 51, Scan time (ns): "52351290" Process id 25, Scan time (ns): "1256822" Process id 86, Scan time (ns): "45630" Process id 9, Scan time (ns): "34561"
解决方案
因为日志是按空行分隔的进程条目,用awk处理最合适,能一次性提取进程ID和目标字段,再排序输出。
1. 按扫描时间(Scan time)降序排序
执行以下命令:
awk -v RS="" '{ pid = ""; time_val = "" for(i=1; i<=NF; i++) { if($i == "Process" && $(i+1) == "id:") pid = $(i+2); if($i == "Scan" && $(i+1) == "time" && $(i+2) == "(ns):") { gsub(/"/, "", $(i+3)); time_val = $(i+3); } } print time_val, "Process id:", pid, "Scan time (ns):", "\"" time_val "\"" }' logfile | sort -nr | cut -d' ' -f2-
命令说明:
-v RS="":将空行设为记录分隔符,每个进程条目作为独立记录处理- 遍历字段提取进程ID和扫描时间(去除引号转为纯数字,方便排序)
- 先输出数字值(用于排序),再输出目标格式内容
sort -nr:按数字从大到小排序cut -d' ' -f2-:去掉开头用于排序的数字,只保留需要的关联内容
2. 按扫描文件总数(Total files scanned)降序排序
把扫描时间的逻辑替换为文件总数即可:
awk -v RS="" '{ pid = ""; file_val = "" for(i=1; i<=NF; i++) { if($i == "Process" && $(i+1) == "id:") pid = $(i+2); if($i == "Total" && $(i+1) == "files" && $(i+2) == "scanned:") file_val = $(i+3); } print file_val, "Process id:", pid, "Total files scanned:", file_val }' logfile | sort -nr | cut -d' ' -f2-
示例输出(按扫描文件数排序):
Process id: 25, Total files scanned: 42 Process id: 86, Total files scanned: 2 Process id: 7, Total files scanned: 0
内容的提问来源于stack exchange,提问作者onthenile
相关产品推荐
相关产品推荐

