Splunk查询问题:匹配相同IP的Tunnel上下线日志并计算会话时长
解决Splunk中Tunnel01上下线日志IP不匹配的会话统计问题
核心问题分析
你当前的transaction仅按host分组,导致同一host下不同IP的上下线日志被错误配对。必须把邻居IP作为额外的分组维度,才能确保同IP的down/up日志对应。
方案1:改进Transaction查询(适合新手理解)
先从原始日志中提取邻居IP,再按host+neighbor_ip分组执行transaction,自动计算中断时长:
index=* "(Tunnel01) is down" OR "(Tunnel01) is up" host=* # 从_raw字段提取Neighbor后面的IP,生成neighbor_ip字段 | rex field=_raw "Neighbor (?<neighbor_ip>\d+\.\d+\.\d+\.\d+) \(Tunnel01\)" # 按host和neighbor_ip分组,匹配同IP的down/up日志 | transaction host neighbor_ip startswith="(Tunnel01) is down" endswith="(Tunnel01) is up" # 保留关键字段,duration是transaction自动计算的中断时长(秒) | table host neighbor_ip _time duration _raw
- 说明:
transaction会自动生成duration字段,值为up日志时间减去down日志时间,单位是秒。 - 可选过滤:如果要排除未完成的会话(只有down没有up,或反之),可以加
| where isnotnull(duration)
方案2:用Stats实现(性能更优,适合大数据量)
transaction在数据量大时性能较差,推荐用stats分组统计,手动计算中断时长:
index=* "(Tunnel01) is down" OR "(Tunnel01) is up" host=* | rex field=_raw "Neighbor (?<neighbor_ip>\d+\.\d+\.\d+\.\d+) \(Tunnel01\) is (?<status>down|up)" # 按host、neighbor_ip分组,分别取最早的down时间和最晚的up时间 | stats earliest(eval(if(status="down", _time, null()))) as down_time latest(eval(if(status="up", _time, null()))) as up_time by host neighbor_ip # 计算中断时长(秒),过滤掉无效会话 | where isnotnull(down_time) AND isnotnull(up_time) | eval duration=up_time - down_time # 转成易读的时间格式(可选) | eval down_time=strftime(down_time, "%Y-%m-%d %H:%M:%S"), up_time=strftime(up_time, "%Y-%m-%d %H:%M:%S"), duration=tostring(duration, "duration") | table host neighbor_ip down_time up_time duration
- 说明:
tostring(duration, "duration")会把秒数转成HH:MM:SS的格式,更易读。
关键注意点
- 确保
rex正则匹配准确:如果日志格式有变化,需要调整正则(比如Tunnel后面的数字可能变化,可以改成\(Tunnel\d+\))。 - 如果
source_ip字段已经包含该邻居IP,可以直接用source_ip代替提取的neighbor_ip,省去rex步骤。
内容的提问来源于stack exchange,提问作者user24042478
相关产品推荐
相关产品推荐

